Skip to content
All library documents

Temporal-Difference Learning with a Separate Policy Network for Trading

Article MQL5 articles

Summary

The article explains temporal-difference (TD) learning as an incremental reinforcement-learning method that updates a state’s value using the immediate reward and the estimated value of the next state. It contrasts this state-value approach with SARSA’s on-policy action values and Q-learning’s off-policy action values, and describes corresponding updates in an MQL5 trading system.

Because a state-value estimate does not itself specify which action to take, the example pairs TD with a neural-network policy classifier. The described three-layer network selects among buy, sell, and hold, while an epsilon-greedy mechanism supports exploration and exploitation. The article outlines the network inputs and discusses configurable architecture and learning settings.

This is an implementation-oriented explanation, not evidence of trading performance: no backtest results or measured returns are provided. The author notes that exploration decay and a variable epsilon could be explored, and that network settings can materially affect behavior. The approach is presented as customizable example code, not a validated trading edge.

Key ideas

  • TD updates state values incrementally from reward and the next state’s estimated value.
  • SARSA uses the next action selected by the current policy, while Q-learning uses the best estimated next action.
  • A separate policy network is used to choose among buy, sell, and hold because state values alone do not define an action.
  • The example combines an epsilon-greedy action choice with a three-layer classifier.
  • The article provides implementation details but no empirical evidence that the system is profitable.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.