Go-Explore for Reinforcement Learning in Trading Strategy Development
Summary
The article introduces Go-Explore as a reinforcement learning approach for problems with sparse rewards and difficult exploration. Rather than immediately optimizing a policy, it first explores and archives visited states and the action paths used to reach them. It then revisits selected archived states, explores onward, and accumulates examples for later policy learning. The MQL5 implementation divides this workflow among separate expert advisors: one explores and records states, another uses the archive, and a later stage fine-tunes a trading policy. Multiple agents are run through the Strategy Tester optimizer, with their archives combined.
The article describes static state and action buffers and file-based archive handling, alongside neural-network components in later stages. It reports that historical-data testing produced strong results, but gives little detail in the supplied passage and acknowledges that the test covered a short period. It therefore calls for longer, more representative training and testing before real-account use. The approach is a method for exploration and policy learning; the text does not establish a durable trading edge.
Key ideas
- Go-Explore prioritizes discovering and revisiting promising states before directly optimizing a policy.
- The exploration phase stores state descriptions and the actions needed to reach them in an archive.
- Revisiting archived states lets agents explore new paths from known points, including in sparse-reward settings.
- The MQL5 design separates exploration, archive-based learning, and policy fine-tuning into distinct expert advisors.
- The reported historical tests cover a short period, so they do not establish robustness in live trading.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.