Skip to content
All library documents

Soft Actor-Critic Optimization Through Policy-Based Action Sampling

Article MQL5 articles

Summary

This article addresses a Soft Actor-Critic trading model that had not trained profitably in an earlier installment. Its central change replaces action perturbations chosen independently of the policy with stochastic sampling from the actor’s learned action distribution. The implementation selects an action quantile using cumulative probabilities and a random draw, then calculates the entropy contribution using the selected action’s probability. Random values are generated by the main program and passed to the OpenCL kernel for sampling during training.

The author argues that uniform perturbations can conflict with SAC’s entropy objective: they may overemphasize actions outside the learned distribution or assign too much exploration to poorly performing actions. The article reports that the changes improved model profitability, while acknowledging that the resulting policy’s optimality is unclear and that architecture, hyperparameters, and training data remain avenues to explore. No detailed performance metrics or evidence about robustness across markets are included in the provided text.

Key ideas

  • SAC combines stochastic policies with an entropy term to balance exploration and exploitation.
  • Sampling actions independently of the actor’s learned distribution can distort the policy’s entropy signal.
  • The proposed implementation samples quantiles according to cumulative learned probabilities.
  • Random draws are created in the main program and sent to OpenCL for use by the training kernel.
  • The author reports improved profitability but provides no detailed metrics establishing the model’s quality or generality.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.