ADX and CCI Pattern Trading with a TRPO Reinforcement Learning Policy
Summary
This article develops a reinforcement learning extension to an Expert Advisor that uses ADX and CCI patterns. It contrasts a more continuous feature representation, where indicator conditions are encoded separately, with a compact representation that combines conditions into bullish or bearish signals. The reported comparison found that only three of ten tested patterns forward-walked successfully from 2024 through 2025 after training on earlier EUR/USD daily data; the authors suggest the longer horizon may favor discrete inputs but leave alternative encodings open for experimentation.
The article explains how actions and rewards extend a supervised model’s forecasts: actions confirm a long or short forecast, while rewards measure trade outcomes and could include favorable and adverse excursions as well as net gain. It introduces TRPO’s policy and value networks, advantage estimates, and KL-divergence constraint as mechanisms for limiting policy changes during training. The excerpt does not provide a full account of final performance or enough evaluation detail to establish robustness; the author presents the RL stage as a cautious improvement and leaves inference work for readers.
Key ideas
- ADX strength and CCI threshold conditions can be represented either as separate features or combined directional patterns.
- The article reports that three of ten patterns forward-walked on EUR/USD daily data under its stated train and test periods.
- The proposed RL actions confirm the supervised model’s directional forecast rather than selecting among order types.
- Rewards can account for profit and loss as well as favorable and adverse trade excursions.
- TRPO constrains policy updates with a KL-divergence trust region to reduce destabilizing changes.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.