Distance-Weighted Supervised Learning for Offline Trading Policies
Summary
The article explains Distance Weighted Supervised Learning (DWSL), an offline method for learning goal-conditioned policies from state-action trajectories that may be suboptimal and lack explicit goal labels. It learns a distribution of distances between states observed in the dataset, then uses a soft minimum estimate of distance to guide the policy toward shorter paths. Action preferences are weighted by the improvement in estimated distance, rather than raw distance, so states near a goal do not automatically dominate training.
The MQL5 implementation adapts the method to a trading agent: it uses an Actor-Critic setup and adds trajectory weighting, while omitting per-step subgoals and a goal-setting model. The author explicitly notes that this implementation differs from the original research method, which was developed for goal-directed robotics and assumes a deterministic environment in its formal setup. The article reports a profitable balance curve in its own testing and generalization to new data, but supplies no basis here for inferring robust live performance. It cautions that the programs demonstrate a technique and are not ready for real trading.
Key ideas
- DWSL aims to learn from offline trajectories without labeled subgoals or known reward labels.
- A classifier estimates distances between dataset states, and a soft minimum reduces sensitivity to approximation errors.
- Actions are weighted by estimated distance reduction so training favors progress toward a goal.
- The MQL5 trading adaptation combines the approach with Actor-Critic training and trajectory weighting.
- The reported test result is a demonstration and does not establish live trading readiness.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.