Skip to content
All library documents

Online Decision Transformers for Fine-Tuning Trading Policies

Article MQL5 articles

Summary

The article explains Online Decision Transformer (ODT), an approach that fine-tunes a Decision Transformer through ongoing interaction with an environment. It reviews how the base model conditions actions on recent states, actions, and return-to-go values, then describes ODT’s stochastic policy, entropy constraint for exploration, and trajectory-based replay buffer. Training samples fixed-length subsequences from stored trajectories, and the initial return-to-go target is set using a scale based on expert results.

The practical section adapts previously trained models for additional online training in an MQL5 trading setup. The implementation retains the earlier architecture and replay data, but omits the explicit entropy term and the proposed initialization with only the highest-return offline trajectories. The article reports improved model efficiency during its online training experiment, while cautioning that the programs demonstrate the technology and are not ready for live trading. It gives no detailed metrics or independent validation in the supplied text, so the reported improvement does not establish generalization or profitability.

Key ideas

  • ODT extends offline Decision Transformer training with further learning from online interaction.
  • It uses a stochastic policy and an entropy lower bound to encourage exploration.
  • Its replay buffer stores trajectories, from which fixed-length subsequences are sampled for training.
  • The described MQL5 implementation omits the entropy loss and uses the full existing buffer rather than seeding it only with top-performing trajectories.
  • The article reports improved model efficiency but gives no detailed validation here and warns against direct live use.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.