Skip to content
All library documents

Reinforcement Learning Exploration with Ensemble Disagreement

Article MQL5 articles

Summary

This article explains disagreement-based exploration in reinforcement learning. An ensemble of forward dynamics models predicts the next environment state from the current state and action. The variance among their predictions becomes an intrinsic reward: actions associated with greater disagreement encourage the agent to explore less familiar regions. As models receive more experience in a region, their predictions should become more consistent. The article distinguishes this approach from rewarding prediction error, which can keep an agent interested in inherently stochastic outcomes.

For its MQL5 adaptation, the author predicts compressed hidden states with an ensemble and evaluates candidate actions in parallel. The described implementation uses fully connected models to keep computation manageable and adds disagreement to predicted reward when selecting actions, while policy training still targets external rewards. The author reports profit in a MetaTrader 5 strategy test, but says training and testing covered a short period and that longer historical training is needed before real trading use. That result is preliminary and does not establish robustness or generalization.

Key ideas

  • Disagreement among forward-model predictions can serve as an intrinsic reward for exploring unfamiliar states.
  • The method uses ensemble output variance rather than prediction error to reduce persistent attraction to stochastic outcomes.
  • The MQL5 adaptation predicts compressed states and compares model predictions across possible actions.
  • In the described setup, disagreement guides action selection while policy training continues to maximize external rewards.
  • The reported strategy test covered a short period, so further historical evaluation is needed before practical use.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.