Evolution Strategies for Neural Network Optimization Without Gradients
Summary
The article explains an evolution strategy for tuning neural network parameters when analytical gradients are unavailable or difficult to calculate. It contrasts the method with genetic algorithms, which select and recombine varied individuals, and estimates a useful search direction by adding random noise to copies of one model. Each perturbed model is evaluated on a training set, and its reward weights the corresponding noise in an update to the original parameters. This treats the model as a whole rather than estimating each parameter’s effect separately.
The article describes an MQL5 implementation for reinforcement learning, including probabilistic action selection and mutation additions that differ from the referenced original approach. It reports that optimization and strategy-tester experiments generated profit, but the test covered only a short interval, so long-term profitability is unknown. The model and implementation are presented as demonstrations that need further configuration and optimization before real-account use.
Key ideas
- Evolution strategies approximate a parameter update direction by evaluating models with small random perturbations.
- The method weights perturbations by each model’s reward before updating the original model parameters.
- Evaluating a perturbed population avoids calculating a separate experimental derivative for every parameter.
- The described MQL5 implementation modifies the referenced approach with probabilistic action selection and mutation.
- Short-interval tester results do not establish long-term profitability.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.