EDL Hierarchical Reinforcement Learning for State-Covering Skills
Summary
Explore, Discover and Learn (EDL) is an unsupervised skill-learning method for hierarchical reinforcement learning. It addresses limited state coverage in approaches that train skills from a prior distribution. EDL first gathers a broad sample of environment states, then uses a variational autoencoder to infer a probabilistic relationship between states and latent skills. Finally, reinforcement learning trains an agent to realize the states associated with each skill, using a fixed reward based on state prediction. This reverses the workflow of training skills first and selecting among them later.
The article describes an MQL5 implementation for a trading agent, including the use of the autoencoder’s latent representation as the agent input to reduce redundant processing. It discusses testing and identifies the quality of state prediction as a key dependency and bottleneck. The method is presented as a way to organize agent behavior, not as a standalone trading signal; market stochasticity and risk remain substantial, and the supplied excerpt gives no specific performance figures to establish profitability or generalization.
Key ideas
- EDL discovers latent skills from sampled environment states before training policies to use those skills.
- A variational autoencoder models a probabilistic mapping between environment states and skill representations.
- The agent is trained with reinforcement learning and a fixed reward derived from predicted states.
- Using the autoencoder’s latent representation alone can reduce duplicated input processing when scheduler and agent share source data.
- The approach depends heavily on accurate state prediction and does not by itself demonstrate profitable trading.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.