Parallelizing Neural Network Training with OpenCL
Summary
The article explains how to speed up neural network training in an MQL5 Expert Advisor by moving repeated neuron calculations to OpenCL kernels. It contrasts the terminal’s thread allocation with OpenCL’s ability to run work across supported CPUs and GPUs, and argues that neurons within a layer can be calculated in parallel while layers themselves remain sequential.
The implementation covers kernels for feed-forward calculations, output and hidden-layer gradients, and weight updates, alongside integration with the network classes. A reported CPU comparison on the same laptop and network architecture found 75 epochs completed in 5 hours and 27 minutes with OpenCL, averaging 4 minutes and 22 seconds per epoch versus 40 minutes and 48 seconds without it. The author reports a 9.35-fold speedup. This is a single hardware and workload example; the article anticipates further gains from a compatible GPU but does not provide GPU benchmark results.
Key ideas
- Neurons in the same layer can be computed in parallel because their calculations do not depend on one another.
- OpenCL kernels perform feed-forward calculations, gradient calculations, and weight updates.
- The implementation keeps layer sequencing intact while parallelizing work across neurons.
- A CPU test reported faster training with OpenCL on the stated laptop and network setup.
- The reported comparison is specific to one workload and does not establish performance on other devices.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.