Skip to content
All library documents

Deep Reinforcement Learning with Elicitable Dynamic Spectral Risk Measures

Article arXiv papers · Author: Anthony Coache et al.

Summary

The document presents a risk-sensitive reinforcement learning framework in which an agent optimizes time-consistent dynamic spectral risk measures. It uses conditional elicitability to construct strictly consistent scoring functions, which act as penalties during estimation. Deep neural networks estimate the risk measures, and the authors prove that this class of measures can be approximated to arbitrary accuracy with such networks.

The framework also includes an actor-critic algorithm that learns from full episodes without adding nested transitions. The authors compare it with a nested simulation approach in two applications: statistical arbitrage and portfolio allocation, using both simulated and real data. The description provides no numerical results, implementation details, or information about the real datasets, so it supports understanding the design and stated evaluation scope but not judging performance gains or robustness. The method is relevant to portfolio decisions where controlling the distribution of losses matters alongside expected reward.

Key ideas

  • The framework optimizes time-consistent dynamic spectral risk measures in reinforcement learning.
  • Conditional elicitability supplies strictly consistent scoring functions used as estimation penalties.
  • Deep neural networks estimate the risk measures, with an approximation result stated for the measure class.
  • The actor-critic method uses full episodes and avoids additional nested transitions.
  • The comparison covers nested simulation, statistical arbitrage, and portfolio allocation on simulated and real data.

Tags

Full text
# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning


# Conditionally Elicitable Dynamic Risk Measures for Deep Reinforcement Learning









We propose a novel framework to solve risk-sensitive reinforcement learning (RL) problems where the agent optimises time-consistent dynamic spectral risk measures. Based on the notion of conditional elicitability, our methodology constructs (strictly consistent) scoring functions that are used as penalizers in the estimation procedure. Our contribution is threefold: we (i) devise an efficient approach to estimate a class of dynamic spectral risk measures with deep neural networks, (ii) prove that these dynamic spectral risk measures may be approximated to any arbitrary accuracy using deep neural networks, and (iii) develop a risk-sensitive actor-critic algorithm that uses full episodes and does not require any additional nested transitions. We compare our conceptually improved reinforcement learning algorithm with the nested simulation approach and illustrate its performance in two settings: statistical arbitrage and portfolio allocation on both simulated and real data.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.