Skip to content
All library documents

Offline Preference-Guided Policy Optimization for Trading Agents

Article MQL5 articles

Summary

The article explains OPPO, an offline reinforcement learning approach that learns a policy directly from comparisons between trajectories. It contrasts this with the common two-stage method of first fitting a scalar reward model and then training a policy against that reward. OPPO instead learns a contextual policy and a preference representation together, aiming to retain richer information about what makes one trajectory preferable. The approach uses paired preference labels and alternates updates to the preference encoder and an optimal context embedding.

The practical example adapts the method for trading in MQL5. It stores context embeddings with states and assigns preference at the trajectory level; the example uses ending profit as its criterion, while noting that drawdown or other measures could also be considered. Separate attention-based models handle preference scheduling and agent behavior. The article reports profitable results on both training and held-out historical periods, but gives limited detail about evaluation methodology in the supplied text. It also notes deviations from the original algorithm and cautions that the programs are demonstrations, not ready for live trading.

Key ideas

  • OPPO learns a policy directly from offline trajectory preferences instead of relying on a separately trained scalar reward model.
  • A high-dimensional preference context is used to represent task information for conditioning the policy.
  • Preference labels compare whole trajectories, and the example uses trajectory profit as its sole criterion.
  • The MQL5 implementation divides preference modeling and behavior learning between two attention-based models.
  • Reported training and test-period profits are not evidence of live-trading readiness, and the implementation differs from the original method.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.