Skip to content
All library documents

DADS: Learning Predictable and Diverse Reinforcement Learning Skills

Article MQL5 articles

Summary

The article explains Dynamics-Aware Discovery of Skills (DADS), an unsupervised reinforcement learning method intended to learn reusable behaviors that are both diverse and predictable. It contrasts DADS with DIAYN: DIAYN encourages distinguishable behavior, while DADS also trains a dynamics discriminator to predict the next state from the current state and selected skill. The agent then receives reward for producing outcomes that fit its skill and differ from outcomes expected across skills.

Training alternates between updating the discriminator and the skill policy using separate batches sampled from a shared experience buffer. The article describes importance weighting to account for differences between the current policy and the policy that generated stored experience. It also outlines an MQL5 implementation with an agent, discriminator, and scheduler, and reports favorable testing results, including a stated rise in average profitable trades to 53.55% and average profitable trade value of 3.08. These results are presented as demonstrations, not proof of robustness: the article says the programs are not ready for live trading, and its focus is skill scheduling rather than long-horizon planning.

Key ideas

  • DADS trains skills to produce outcomes that are both distinguishable and predictable.
  • A discriminator predicts the next state given the current state and the selected skill.
  • The skill policy is rewarded for matching its predicted outcome while differing from the average across skills.
  • Training alternates between the dynamics model and skill policy using samples from a shared experience buffer.
  • The article reports test results but cautions that its implementation is demonstrative and not ready for live trading.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.