Reward-Free Pretraining for Decision Transformers with Future Trajectory Embeddings
Summary
The article explains Pretrained Decision Transformer (PDT), a reinforcement learning approach that first trains on offline trajectories without reward labels and later adapts to a reward-defined task. During pretraining, an actor predicts actions from past states and actions while a future encoder compresses a segment of upcoming trajectory into a latent embedding. A target predictor learns to infer that embedding from the current state, creating a way to condition behavior on learned future patterns without using returns as labels.
For downstream learning, a reward prediction model associates future embeddings with task outcomes, allowing the model to favor embeddings linked to higher rewards. The article also describes an MQL5 implementation and dataset collection based on previously gathered experience, including exploratory behavior. It presents the method conceptually and discusses practical training architecture, but the supplied excerpt does not provide quantitative trading results or a comparative evaluation. The authors note that training is computationally demanding, dataset needs can be substantial, and the tradeoff between behavior diversity and consistency depends on the data.
Key ideas
- PDT pretrains a Decision Transformer on trajectories without reward labels.
- A future encoder compresses upcoming trajectory information into a latent embedding used to condition action prediction.
- A separate predictor estimates future embeddings from the current state, supporting reward-free behavior learning.
- Fine-tuning uses reward prediction to associate embeddings with task outcomes and guide behavior.
- Compute requirements, dataset size, and the balance between diverse and consistent behavior remain practical constraints.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.