Skip to content
All library documents

Using DDPG to Refine Moving-Average and Stochastic Trading Signals

Article MQL5 articles

Summary

The article introduces Deep Deterministic Policy Gradient (DDPG) as a reinforcement-learning filter for signals derived from moving-average and stochastic-oscillator patterns. It contrasts DDPG’s continuous-valued action output with discrete buy, sell, or hold classification: a value around a central threshold can map to directional choices, while a wider action range could support options such as position sizing. The proposed agent uses an actor to select actions and a critic to estimate their value, with target networks, experience replay, and action noise to support training and exploration.

The discussion explains replay-buffer sampling, Bellman-based critic updates, and gradual target-network updates. The strategy results are not presented here: the author defers the agent, environment, and tester reports to a later article. The preceding work had only a one-year test window, and the author notes that few of the tested signal patterns traded both directions, urging evaluation on more history. The described method is therefore an implementation proposal, not evidence of a validated trading edge.

Key ideas

  • DDPG learns a continuous action policy using an actor network and a critic that estimates action value.
  • A replay buffer stores past state-action transitions and randomly samples them to reduce correlation between training updates.
  • Slowly updated target networks are used to stabilize learning targets.
  • The proposed reinforcement learner is intended to filter signals formed from moving averages and a stochastic oscillator.
  • The article does not report DDPG strategy results and notes that earlier pattern tests used a limited history window.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.