Skip to content
All library documents

Reinforcement Learning and the Cross-Entropy Method for Trading

Article MQL5 articles

Summary

The article introduces reinforcement learning through the interaction of an agent and an environment. The agent observes a state, chooses an action under a policy, and receives rewards as the environment changes. It contrasts this setup with supervised learning, where examples have target answers, and unsupervised learning, where static data is grouped or structured. Trading is framed as a sequential problem because position outcomes may only be clear when a trade closes.

The article emphasizes that training requires a finite evaluation period and a carefully designed reward system, especially when rewards arrive after a sequence of actions. Poor reward attribution can reinforce bad entries or discourage good ones. It then presents the cross-entropy method as a heuristic for searching strategies through repeated trials, with an MQL5 implementation discussed. The approach is described as accessible but limited; the text does not establish robust live trading performance. Results depend on the chosen reward definition, training sample, and how well the environment represents future market conditions.

Key ideas

  • Reinforcement learning trains an agent through actions and feedback from an environment.
  • A trading agent’s objective is to maximize cumulative reward over a bounded evaluation period.
  • Delayed trade outcomes make it difficult to assign reward fairly to entry and exit decisions.
  • Reward design can steer learning toward undesirable behavior if it misrepresents action quality.
  • The cross-entropy method is introduced as a heuristic strategy-search approach with implementation limitations.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.