Skip to content
All library documents

Adaptive Gradient Methods for Neural Network Training with Adam

Article MQL5 articles

Summary

The article surveys adaptive neural network optimizers, explaining how AdaGrad, RMSProp, Adadelta, and Adam adjust parameter updates using accumulated gradient information. It describes AdaGrad’s shrinking learning rate, RMSProp’s exponentially weighted squared gradients, and Adadelta’s use of past update magnitudes. The practical implementation focuses on adding Adam to an existing neural network in MQL5, including an OpenCL kernel that updates weights and stores moving averages of gradients and squared gradients.

In a fractal-classification experiment, the author reports that Adam reached 48.6% correct classification after five training epochs before declining to 41.1%; stochastic gradient descent stayed below 10% after 90 epochs. The author also reports fewer missed fractals with Adam. These figures are from a single described task and do not establish broader predictive or trading performance. The comparison depends on the dataset, model, training settings, and implementation.

Key ideas

  • Adaptive optimizers change learning updates using gradient history to handle parameters with different behavior.
  • AdaGrad reduces updates for frequently changing parameters but its accumulated gradient squares can drive the learning rate toward zero.
  • RMSProp uses an exponential average of squared gradients, while Adam tracks averages of gradients and their squares.
  • The implementation adds Adam weight updates to an MQL5 neural network, including an OpenCL version.
  • The reported fractal-classification results favor Adam over stochastic gradient descent in the tested setup, but do not establish generalization to trading.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.