Skip to content
All library documents

Actor–Director–Critic Reinforcement Learning for More Selective Trading Policies

Article MQL5 articles

Summary

The article describes an Actor–Director–Critic (ADC) deep reinforcement learning design intended to guide policy learning when a conventional Actor–Critic's value estimates are unreliable early in training. The Actor proposes actions, the Critic estimates their long-term value, and a Director classifies actions as suitable or unsuitable using labeled high- and low-performing state–action–reward examples. The Director's influence is intended to decay as the Critic becomes more accurate, reducing its role over time.

To address value overestimation, the framework uses two Critics and pairs each with two target networks. It averages each pair of target estimates for learning targets, uses the lower estimate from the Critics during Actor training, and delays policy updates. These mechanisms are described as ways to reduce variance and stabilize training. The article explains the architecture and its intended benefits, but does not present trading results: evaluation on historical market data is deferred to a future article. Its claims about faster or safer learning should therefore be treated as proposed design motivations, not demonstrated trading performance.

Key ideas

  • The Director classifies candidate actions to guide the Actor while Critic estimates are still unreliable.
  • The Director is trained on labeled examples of higher- and lower-performing actions.
  • The Director's contribution is reduced as the Critic improves.
  • Paired target networks and averaged estimates are proposed to reduce value-estimation instability.
  • The article presents architecture details but defers evaluation on historical market data.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.