Using a Hierarchical Decision Transformer for Goal-Directed Trading
Summary
The article adapts ideas from Control Transformer, a hierarchical reinforcement learning method designed for robot navigation, to trading. The approach separates a low-level policy that handles local transitions from a higher-level sequence model that plans toward a global goal. In the original navigation setup, a probabilistic road map supplies waypoints and goal-conditioned learning trains the local controller. A value model estimates expected reward for a state and goal, providing a starting return-to-go estimate that can be adjusted using realized rewards.
For the trading implementation, the article describes collecting candidate trajectories through uniform random action sampling and selecting profitable paths, then dividing model training among separate Expert Advisors for scheduling, local policy learning, and value estimation. It reports that testing produced a model capable of generating profit, but gives no detailed metrics or enough evidence to assess robustness. The author cautions that the programs are demonstrations, not intended for real-market use, and notes that offline policies can suffer from distribution shift; policy-generated data and subsequent offline fine-tuning are proposed as a mitigation.
Key ideas
- The method divides long-horizon control into a local policy and a higher-level planner that pursues a global goal.
- A value model estimates expected reward conditioned on the current state and goal to initialize return-to-go.
- The trading example gathers candidate trajectories using uniformly sampled actions and selects profitable sequences for training.
- Separate Expert Advisors handle distinct training tasks, enabling parallel model development.
- The article claims profitable test behavior but supplies no detailed performance evidence and warns against real-market use.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.