Skip to content
All library documents

Deploying TD3 Trading Signals with Ichimoku and ADX Features

Article MQL5 articles

Summary

This article describes a workflow for training a Twin Delayed Deep Deterministic Policy Gradient (TD3) agent in Python and deploying its actor network in an MQL5 Expert Advisor through ONNX. It presents the trading task as sequential decision-making: historical price features, including Ichimoku and ADX measures, form the state; the model emits a continuous action; and rewards account for market movement and transaction costs. The discussion explains TD3’s twin critics, smoothed target actions, and delayed actor updates as ways to address value overestimation and unstable learning in continuous-action problems.

The author outlines replay-buffer training and relevant hyperparameters, then reviews forward tests on three signal patterns. Results are mixed: only one pattern is described as modestly encouraging. The article also acknowledges that its use of a fixed exported actor during inference departs from approaches that continue learning in deployment, and that feature preprocessing may be suboptimal. These Strategy Tester results are preliminary; they do not demonstrate live profitability or broad robustness across market conditions.

Key ideas

  • TD3 is presented as a continuous-action reinforcement-learning method for trading decisions such as position sizing.
  • Twin critics, target-policy smoothing, and delayed policy updates are intended to improve on DDPG stability.
  • The workflow trains in Python and exports the actor network for inference in an MQL5 Expert Advisor.
  • The example builds states from price-derived features, including Ichimoku and ADX measures, and includes transaction costs in rewards.
  • Forward tests across three signal patterns are mixed, with only one showing modestly encouraging results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.