Temporal-Difference Learning with a Separate Policy Network for Trading
Summary
The article explains temporal-difference (TD) learning as an incremental reinforcement-learning method that updates a state’s value using the immediate reward and the estimated value of the next state. It contrasts this state-value approach with SARSA’s on-policy action values and Q-learning’s off-policy action values, and describes corresponding updates in an MQL5 trading system.
Because a state-value estimate does not itself specify which action to take, the example pairs TD with a neural-network policy classifier. The described three-layer network selects among buy, sell, and hold, while an epsilon-greedy mechanism supports exploration and exploitation. The article outlines the network inputs and discusses configurable architecture and learning settings.
This is an implementation-oriented explanation, not evidence of trading performance: no backtest results or measured returns are provided. The author notes that exploration decay and a variable epsilon could be explored, and that network settings can materially affect behavior. The approach is presented as customizable example code, not a validated trading edge.
Key ideas
- TD updates state values incrementally from reward and the next state’s estimated value.
- SARSA uses the next action selected by the current policy, while Q-learning uses the best estimated next action.
- A separate policy network is used to choose among buy, sell, and hold because state values alone do not define an action.
- The example combines an epsilon-greedy action choice with a three-layer classifier.
- The article provides implementation details but no empirical evidence that the system is profitable.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.