Skip to content
All library documents

Implementing a DDPG Agent for Trading with Indicator-Based States

Article MQL5 articles

Summary

The article explains a deep deterministic policy gradient agent for a trading workflow that combines supervised-learning outputs with reinforcement learning. Its agent uses actor and critic networks, matching target networks, separate optimizers, and a replay buffer. The policy produces continuous actions, with Gaussian noise added for exploration and action values clipped to the allowed range.

Learning updates estimate critic targets with the Bellman relation, train the critic using mean squared error, update the actor to favor higher critic values, and softly track target-network weights. The article also discusses saving model states, retrieving MetaTrader 5 price data in Python, and using moving-average and stochastic patterns as inputs. It reports that only three of seven models forward-walked on the subsequent year, with trades often concentrated in one direction. The author cautions that the test window was small; the results are limited and do not establish robust performance.

Key ideas

  • The DDPG agent separates policy learning in an actor from value estimation in a critic.
  • Target networks and replay-buffer sampling are used to stabilize off-policy updates.
  • Gaussian action noise supports exploration for the deterministic continuous-action policy.
  • Moving-average and stochastic patterns contribute to the trading state representation.
  • The reported forward walk was mixed, and the article attributes possible distortions to a short test window.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.