Skip to content
All library documents

Decomposing Reinforcement Learning Rewards for Trading Agents

Article MQL5 articles

Summary

This article describes reward decomposition for reinforcement learning, with an implementation in a Soft Actor-Critic variant using the DICE method. Rather than learning one value function from a composite reward, the agent learns a value estimate for each reward component; a weighted combination then supports policy optimization and comparison of different reward priorities. The components can also help diagnose which parts of the reward signal influence behavior.

The discussion highlights implementation challenges: training multiple value estimates increases complexity, while independently taking component-wise minimum critic estimates can create imbalance. The described approach instead selects the critic with the lower overall score and uses its component estimates. It also applies Conflict-Averse Gradient Descent to address competing objectives during optimization. The article reports that its implementation produced profit on both training and out-of-sample data, but calls the results insufficient and encourages experiments with reward components. This is a reported result, not evidence of robust performance across markets or settings.

Key ideas

  • Reward decomposition trains a separate value estimate for each component of a composite reward.
  • A weighted sum of component values can reconstruct the overall objective and evaluate different priorities.
  • Component-wise minimum critic estimates may imbalance training; selecting a critic by its total estimate is proposed instead.
  • Conflict-Averse Gradient Descent combines component gradients to manage competing optimization objectives.
  • The reported SAC+DICE results were profitable in and out of sample, but the article describes them as preliminary.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.