Skip to content
All library documents

Training a Stochastic Policy Gradient Trading Agent

Article MQL5 articles

Summary

The article explains a reinforcement learning approach that trains a neural network to choose trading actions directly. Unlike Deep Q-Learning, which estimates expected rewards for actions, the policy model outputs a probability distribution over actions. Sampling from that distribution adds behavioral variation, while training shifts probability toward actions associated with positive rewards. A SoftMax layer converts model outputs into normalized action probabilities.

For learning, the agent records states, actions, and rewards during a trading session, then uses accumulated rewards and a classification-style loss to update the policy. The article describes an MQL5 implementation and reports Strategy Tester results, including a version that reached 60% profitable operations after changing the reward policy and model settings. It also reports an average holding time of 1 hour 40 minutes. These results are specific to the author’s tests; they do not establish robustness across market conditions. The article emphasizes that reward design and loss-function choice affect outcomes, and provides no broad out-of-sample evidence.

Key ideas

  • A policy gradient model predicts probabilities for available actions instead of estimating their expected rewards.
  • SoftMax normalizes the model output into a probability distribution over actions.
  • Sampling actions from the distribution supports exploration, while training increases probability for actions linked to positive rewards.
  • The agent stores states, actions, and rewards across a session before updating the policy model.
  • The reported Strategy Tester results depend on the reward policy and model settings and do not demonstrate general performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.