Skip to content
All library documents

Soft Actor-Critic for Continuous-Action Reinforcement Learning

Article MQL5 articles

Summary

The article introduces Soft Actor-Critic (SAC), an off-policy reinforcement learning method for continuous action spaces. Its stochastic actor learns a probability distribution over actions, while an entropy term in the objective balances reward seeking with exploration. Two critics estimate action value, and the lower estimate is used to reduce overestimation. Unlike the described TD3 approach, SAC uses the current actor for action selection and updates the actor and target models at every training step.

The article also describes an MQL5 implementation that builds on a fully parameterized quantile network and adds components for entropy parameters and log probabilities. It discusses replay-buffer training and the choice or adaptation of the entropy temperature. The author reports that the attempted training did not produce a profitable strategy, although the model was said to perform stably on and beyond the training set. That result is not quantified, and the article treats finding a profitable policy as unresolved; stable performance alone does not demonstrate trading profitability or robustness.

Key ideas

  • SAC maximizes expected reward while rewarding policy entropy to preserve action diversity.
  • A stochastic actor represents a distribution over continuous actions rather than selecting one fixed action per state.
  • Two critics are trained, and the lower estimated value informs the target and actor updates.
  • The described procedure uses replay data and updates the actor and target models at each training step.
  • The reported implementation did not yield a profitable strategy, leaving practical trading efficacy unresolved.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.