Evaluating a Gold Trading Bot with Signal Tests and Walk-Forward Validation
Summary
This article lays out a practical pipeline for deciding whether to develop and evaluate a reinforcement-learning trading system for gold. It begins with a supervised LightGBM baseline using triple-barrier labels and purged walk-forward folds to check whether the input features contain directional information. The author reports near-coin-flip AUC results across several horizons and feature sets for M15 XAUUSD data, which led them to question whether the features, rather than the model architecture, contained useful signal.
The rest of the pipeline addresses common sources of misleading results: overfitting to historical episodes, changing market regimes, reward designs that encourage undesirable behavior, and leakage from random splits. It recommends time-ordered validation with purge and embargo periods, promotion checks across multiple seeds, and judging candidates by realized equity and trades rather than shaped rewards. Deployment safeguards include saved normalization, a manifest contract, warm-up requirements, and broker reconciliation. The reported demo outcomes illustrate instability; they do not establish future profitability, and the author calls for more rigorous evaluation before live use.
Key ideas
- A supervised baseline can test for directional information before investing in a more complex reinforcement-learning system.
- Triple-barrier labels and purged walk-forward folds are used to reduce overlap and leakage in time-series validation.
- The author's gold feature tests produced AUC values close to chance across the reported horizons and feature sets.
- Poorly designed rewards can lead agents to oversize positions, avoid trading, or favor high win rates despite negative expectancy.
- Evaluation should focus on realized trades and equity, with additional checks across seeds and deployment safeguards.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.