Scaling Go-Explore Reinforcement Learning to Longer Trading Histories
Summary
The article examines why Go-Explore’s two-stage process—exploring states and then training a policy from collected examples—becomes harder as a trading history grows. It reports that random action selection produced a profitable pass on a one-month training period but none after extending the period to three months. The author links this to the growing state space, the lower chance of a long sequence of successful actions, and possible changes in market conditions.
The proposed adjustments include sizing a fixed path buffer for the chosen history and timeframe, removing an expensive example-sorting step, and changing reward calculation to combine equity and balance changes. It also discusses dimensionality reduction, feature selection, confidence-based exploration, and retraining or model-based methods when conditions shift. The article describes implementation choices and testing on training and test samples, but the supplied text omits much of the implementation and the specific test results. Its broad claims of model quality therefore cannot be independently assessed; it also stresses that changing markets prevent any guarantee of future success.
Key ideas
- Longer training histories make state exploration and successful action sequences more difficult.
- The article reports no profitable pass when random action selection was extended from one month to three months.
- A fixed path buffer should be sized to cover the full training history and instrument trading schedule.
- Removing example sorting may reduce data preparation time as the collected sample grows.
- Combining equity and balance changes is proposed as a reward that accounts for open-position outcomes and realized profit.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.