SPOT Uses Conditional VAE Density Estimates to Constrain Offline RL Policies
Summary
This article explains Supported Policy Optimization (SPOT), an offline reinforcement learning method intended to reduce unreliable value estimates when a learned policy chooses actions that are poorly represented in its fixed training data. SPOT estimates the behavior policy's state conditioned action density with a conditional variational autoencoder (CVAE). A penalty based on estimated log likelihood is added to policy learning, encouraging proposed actions to remain within the data's support. The article describes adapting this idea to a Soft Actor Critic based trading model, with density model training separated from policy training.
The implementation trains the CVAE on state and action pairs, using reconstruction error as an indirect signal of whether proposed actions resemble the dataset. The account describes training and evaluation stages and reports stable learning and a profitable actor behavior in its experiment, but the supplied text gives no detailed performance figures or benchmark comparison. The method can reduce risky extrapolation, but constraining a policy also limits exploration. The author argues it is most appropriate when the offline dataset contains suboptimal trajectories that the learner can improve upon.
Key ideas
- Offline policies can overestimate actions that are poorly represented in their fixed training data.
- SPOT adds a density based penalty to encourage actions supported by the dataset.
- A conditional variational autoencoder estimates action likelihood given the environment state.
- The described implementation trains the density model before training the trading policy.
- Support constraints can improve training stability while limiting exploration beyond the dataset.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.