Skip to content
All library documents

TD3 Improvements for More Stable Continuous-Action Trading Agents

Article MQL5 articles

Summary

The article explains how Twin Delayed Deep Deterministic Policy Gradient (TD3) modifies DDPG to reduce overestimation of learned Q-values. Its three core changes are training two Critics and using the lower estimate, updating the Actor less often, and smoothing target actions. The implementation described uses the first two changes, with target smoothing omitted because the author considers financial market data inherently stochastic.

The practical section adapts the trading robot to maintain positions, add to or partially close them, and manage stop loss and take profit levels. Historical-data experiments are reported as showing profit on both training and new data, with comparable results across the two. However, the article gives no detailed performance statistics in the supplied text, acknowledges that the results still need work, and presents them as a basis for further research rather than evidence of live trading reliability.

Key ideas

  • TD3 reduces Q-value overestimation by training two Critics and using their lower estimate for learning.
  • The method delays Actor updates so the Actor learns from a more stable Critic.
  • Target-action smoothing is a third TD3 feature, but the described trading implementation omits it.
  • The trading robot manages ongoing positions through additions, partial closures, and directional stop loss trailing.
  • The reported historical results include new data, but do not establish live-market performance.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.