SAC+DICE: Optimistic Exploration and Replay Distribution Correction
Summary
The article explains SAC+DICE, an off-policy reinforcement learning approach that combines optimistic exploration with correction for the mismatch between replay-buffer samples and the current policy. It motivates the method by noting that conservative action evaluation can limit exploration, while an exploratory actor can gather useful experience beyond the target policy’s behavior. The approach trains separate exploratory and target actors alongside critics, using upper and lower estimates of action value to guide their roles.
A learned correction ratio, estimated with a DICE-style objective, adjusts learning signals for actors and critics to account for differences in state-action distributions. The article outlines the training sequence and describes an MQL5 implementation that packages the models in a dedicated class. It reports that an EA earned a profit on a test using new data, but gives limited evidence in the supplied text and cautions that the programs are demonstrations, not ready for live markets. Further refinement and testing are needed.
Key ideas
- SAC+DICE uses an optimistic actor to explore while a target actor learns with a conservative lower estimate of action value.
- A DICE-derived correction ratio addresses differences between replay-buffer samples and the current policy’s state-action distribution.
- The method applies distribution correction to actor and critic training objectives.
- The MQL5 implementation coordinates multiple trainable and target models in a dedicated class.
- A profitable test is reported, but the article says the example systems need further testing before live use.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.