Skip to content
All library documents

Goal-Conditioned Reinforcement Learning for Trading Tasks

Article MQL5 articles

Summary

The article explains goal-conditioned reinforcement learning (GCRL) as a way to make an agent’s action depend on both market state and a specified subtask, such as opening or closing a position. It contrasts this with skill conditioning and describes task-specific rewards that supplement the overall trading objective. Such rewards need careful scaling: if they outweigh trading gains and losses, the agent may pursue task completion while damaging account value. Task descriptions can be simple indicators or richer learned representations.

For its MQL5 example, the author combines a state encoder with an agent in a single model. Historical prices and indicators inform the encoder’s market-regime representation, while account information helps define whether the current task concerns entry or exit. The article reports testing outside the training data and notes an unresolved directional bias: the EA opened positions in only one direction, though it still made a profit during the test. This limited result does not establish durable performance, and the reward design and evaluation remain important caveats.

Key ideas

  • GCRL conditions actions on a stated subtask as well as the current environment state.
  • Task-specific rewards can guide entry and exit behavior but must remain balanced against trading outcomes.
  • A task vector may use explicit indicators or a learned representation, depending on the subtask.
  • The example combines an encoder and agent, using market history for regime information and account state for task context.
  • The reported test had a one-direction trading bias, so its positive result does not establish general robustness.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.