Skip to content
All library documents

RE3 Exploration with Fixed Random Encoders and State Entropy

Article MQL5 articles

Summary

The article explains Random Encoders for Efficient Exploration (RE3), a reinforcement learning method for encouraging agents to explore high-dimensional state spaces. It estimates state novelty from distances to nearby states in a low-dimensional representation produced by a randomly initialized, fixed encoder. Those distances feed an entropy-based internal reward, which is combined with the environment’s external reward during policy training. The article also describes storing encoded states in the replay buffer to reduce repeated computation and gradually lowering the exploration reward’s weight as training proceeds.

The practical implementation combines RE3 ideas with an Actor-Critic approach and other previously discussed methods. The author reports that the resulting model had a high share of profitable trades but placed very few trades, so its trading usefulness remains uncertain and requires further work. This is an individual implementation rather than a clean reproduction or controlled comparison of RE3; the excerpt provides no quantitative performance details or evidence isolating the method’s contribution.

Key ideas

  • A fixed, randomly initialized encoder can provide a useful representation for comparing states without training an additional representation model.
  • RE3 estimates state novelty from distances to nearby encoded states and uses that estimate as an internal exploration reward.
  • Saving encoded states in the replay buffer can avoid repeatedly processing high-dimensional observations during training.
  • The exploration reward can be combined with external reward, with its weight reduced over training.
  • The reported trading model had few trades, limiting conclusions about its practical effectiveness.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.