Using SARSA as an On-Policy Trading Signal in MQL5
Summary
The article explains SARSA, an on-policy reinforcement-learning method, and describes using it as an independent signal model in an MQL5 Expert Advisor. Its Q-value update uses the action actually selected in the next state under the agent’s current policy, unlike Q-Learning’s update based on the best possible next action. Epsilon-greedy selection balances exploration and exploitation, though its random choices can make results inconsistent.
The example simplifies the market to nine states built from bullish, bearish, or whipsaw conditions across short and long time frames. The article contrasts this table-based approach with Deep Q-Networks, which use neural networks and can handle larger state spaces. It also discusses state aggregation as a way to represent broader inputs more compactly. These are implementation concepts and qualitative comparisons, not evidence of trading profitability; the author notes that the setup restricts the environment and that larger or continuous state spaces challenge tabular SARSA.
Key ideas
- SARSA updates Q-values using the next action selected by the same policy the agent follows.
- Epsilon-greedy action selection mixes exploration with exploitation, and its randomness can produce inconsistent outcomes.
- The example represents markets with nine states derived from three price-condition categories across two time frames.
- Tabular SARSA is limited by large state spaces, while neural-network methods can generalize across more states.
- Grouping similar conditions into abstract states can make a richer environment more manageable.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.