Dichotomy of Control for Decision Transformer Trading Agents
Summary
The article introduces Dichotomy of Control (DoC), a reinforcement learning approach that separates effects a policy can influence from environmental randomness. It explains the motivation through Decision Transformers: conditioning an agent on a desired return may lead to poor actions when that return is unlikely from the current state. DoC uses a latent representation that should avoid encoding unknown future rewards and states, with a mutual-information constraint intended to separate policy-controlled behavior from stochastic outcomes.
For deployment, the method samples candidate latent states from a state-conditioned prior and ranks them with a value function. The article then describes a personal MQL5 adaptation that combines ideas from fully parameterized quantile functions to simplify latent-state generation and evaluation, reusing autoregressive modeling concepts from the Decision Transformer. It reports a visible improvement on a test sample, but gives no quantitative results in the supplied text. The implementation is explicitly an interpretation that departs from the original method, and the author says the performance remains imperfect.
Key ideas
- DoC distinguishes policy-controlled effects from environmental stochasticity.
- A latent state is intended to represent strategy behavior without encoding unknown future outcomes.
- Candidate latent states can be ranked by estimated value before being passed to a policy.
- The MQL5 implementation adapts and simplifies the original method using quantile modeling ideas.
- Reported test-sample improvement is qualitative, and the implementation is not the original algorithm.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.