Skip to content
All library documents

Distance-Weighted Supervised Learning for Offline Trading Policies

Article MQL5 articles

Summary

The article explains Distance Weighted Supervised Learning (DWSL), an offline method for learning goal-conditioned policies from state-action trajectories that may be suboptimal and lack explicit goal labels. It learns a distribution of distances between states observed in the dataset, then uses a soft minimum estimate of distance to guide the policy toward shorter paths. Action preferences are weighted by the improvement in estimated distance, rather than raw distance, so states near a goal do not automatically dominate training.

The MQL5 implementation adapts the method to a trading agent: it uses an Actor-Critic setup and adds trajectory weighting, while omitting per-step subgoals and a goal-setting model. The author explicitly notes that this implementation differs from the original research method, which was developed for goal-directed robotics and assumes a deterministic environment in its formal setup. The article reports a profitable balance curve in its own testing and generalization to new data, but supplies no basis here for inferring robust live performance. It cautions that the programs demonstrate a technique and are not ready for real trading.

Key ideas

  • DWSL aims to learn from offline trajectories without labeled subgoals or known reward labels.
  • A classifier estimates distances between dataset states, and a soft minimum reduces sensitivity to approximation errors.
  • Actions are weighted by estimated distance reduction so training favors progress toward a goal.
  • The MQL5 trading adaptation combines the approach with Actor-Critic training and trajectory weighting.
  • The reported test result is a demonstration and does not establish live trading readiness.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.