Skip to content
All library documents

Intrinsic Curiosity Rewards for Reinforcement Learning in Trading

Article MQL5 articles

Summary

This article explains the Intrinsic Curiosity Module as a way to encourage reinforcement learning agents to explore when external rewards are sparse. The module encodes consecutive states, learns to infer the action that connected them, and predicts the next encoded state from the current state and action. Prediction error supplies an intrinsic reward, which is combined with the environment reward to guide learning. The article then describes an MQL5 implementation with experience replay and a trading state that includes account and open-position information. Its action set includes buying, selling, closing positions, or waiting.

The cited research reports game-based tests in which curiosity supports exploration, learning from delayed or absent rewards, and transfer to new scenarios. The article adapts the idea to trading, where its external reward is balance change, but does not establish robust trading performance. It explicitly presents the EA as an evaluation example that needs substantial refinement and comprehensive testing before real use.

Key ideas

  • The Intrinsic Curiosity Module rewards prediction error to encourage exploration when external rewards are rare.
  • An inverse model learns actions from pairs of encoded states, while a forward model predicts the next state.
  • The MQL5 example combines curiosity with balance-based rewards and stores experiences in a replay buffer.
  • The article's cited evidence comes from games, and it does not demonstrate reliable live trading results.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.