Decision Transformers for Goal-Conditioned Trading Reinforcement Learning
Summary
The article explains Decision Transformer as an offline reinforcement learning method that frames action selection as autoregressive sequence modeling. Its trajectory representation interleaves return-to-go, state, and prior action tokens; at inference, a desired return and initial state condition action generation. After each environment step, the target return is reduced by the reward received. The described architecture embeds each modality, adds time information, and uses a GPT-style transformer to predict actions from a recent context window.
The implementation discussion adapts this approach to trading data with distinct modalities, including price history and account state, and describes embedding and parallel-processing choices in MQL5 and OpenCL. The reported trading experiment did not remain profitable through the test period: its profit factor was below one and fewer than half of trades were profitable. The author therefore presents the approach as requiring further work. The article gives a method and a concrete negative result, but the supplied text does not establish that the architecture generalizes or will produce profitable strategies.
Key ideas
- Decision Transformer models action sequences conditioned on states, prior actions, and desired return-to-go.
- At inference, the desired return is updated after each reward and used to condition the next action.
- Offline training uses trajectory segments and supervised action prediction.
- Trading inputs may combine price history, account information, reward, actions, and time as separate modalities.
- The reported test did not show sustained profitability, so the implementation remains experimental.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.