Skip to content
All library documents

Distributed Q-Learning for Modeling Reward Distributions

Article MQL5 articles

Summary

The article introduces distributed, or distributional, Q-learning as an extension of classical Q-learning. Instead of predicting one expected reward for each state-action pair, the model estimates a probability distribution over rewards, represented by quantiles within a configured range. A SoftMax operation normalizes the probabilities, and LogLoss replaces the usual scalar prediction loss. The distribution can reveal the chances of both positive and negative outcomes when an average reward would obscure them, supporting risk-aware action selection.

Training retains familiar Q-learning components, including experience replay, a discount factor, the Bellman equation, and a frozen target network. The author recommends first training without future-state targets, then using the trained model as the target network and gradually extending predictions further ahead. The discount factor shapes the balance between near-term and longer-term rewards. An MQL5 implementation and strategy-tester example are described, with the article reporting potential profitability but offering no detailed performance figures in the provided text. The method still depends on chosen reward bounds and quantile count, and testing does not establish reliable live trading performance.

Key ideas

  • Distributional Q-learning estimates a reward distribution rather than a single expected reward.
  • The reward range is represented by quantiles, with bounds and quantile count set as model parameters.
  • SoftMax-normalized probabilities and LogLoss support distribution prediction for each state-action pair.
  • Experience replay, a discount factor, and a target network preserve key elements of classical Q-learning.
  • The article proposes staged target-network training to reduce distortion from initially unreliable future predictions.

Tags

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.