Learning Diverse Trading Skills with Reward-Free Reinforcement Learning
Summary
This article adapts the Diversity Is All You Need approach to training trading agents. Its first phase learns a set of skills without an external task reward: a skill-conditioned agent acts from the observed market and account state, while a discriminator tries to infer the selected skill from the resulting state. Discriminator feedback supplies an intrinsic learning signal intended to make the skills produce distinguishable behavior. The agent is trained with reinforcement learning, while the discriminator uses supervised learning.
In a second phase, a scheduler selects among the learned skills to pursue a task objective, potentially while the skill model is fixed. The implementation discussion describes separate agent, scheduler, and discriminator models, with the discriminator used during training rather than deployment. The article outlines data collection, training, and testing in a MetaTrader workflow, but the supplied excerpt does not provide quantitative performance evidence or establish profitability. The quality of the resulting skills and their usefulness across changing market conditions remain open questions.
Key ideas
- The method first trains a library of distinguishable behaviors without an external task reward.
- A discriminator infers which skill was used from the next state, providing an intrinsic learning signal to the skill model.
- A scheduler is trained afterward to choose skills for a specific objective.
- The discriminator supports skill training but is not required for the deployed agent and scheduler.
- The article describes a trading implementation but does not report evidence that it improves trading performance.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.