Skip to content
All library documents

Training Small Neural Networks with Levenberg–Marquardt

Article MQL5 articles

Summary

The article compares neural-network optimization methods for quickly retraining feed-forward models, with an emphasis on Levenberg–Marquardt (LM) for online adaptation. It explains output-layer error calculation for different loss functions and activation functions, then reviews gradient descent, momentum, and stochastic gradient descent before presenting an LM implementation. Momentum smooths parameter updates and can reduce oscillation; LM uses curvature information to seek a local loss minimum in fewer training passes.

The demonstrations use synthetic data consisting of a periodic signal plus Gaussian noise, rather than market observations. The article reports that LM is fastest among the compared methods for networks with up to roughly 100 hidden neurons, while L-BFGS performs better as network size grows; it also cautions against fitting below the noise-implied loss threshold because that risks overfitting. These results are tied to the stated experiments and do not establish that LM improves trading forecasts or generalizes to real market data.

Key ideas

  • The article implements gradient descent, momentum, stochastic gradient descent, and Levenberg–Marquardt for multilayer perceptrons.
  • Correct output-layer gradients depend on both the chosen loss function and the final activation function.
  • Momentum smooths parameter updates and can improve convergence when ordinary gradient descent oscillates.
  • The reported synthetic-data comparison favors Levenberg–Marquardt for smaller networks, while L-BFGS becomes faster as networks grow.
  • Training below the noise-implied loss threshold can overfit, and the experiments do not demonstrate market forecasting performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.