Skip to content
All library documents

Convex Risk Objectives in Reinforcement Learning for Trading

Article arXiv papers · Author: Shanyu Han et al.

Summary

This paper presents a reinforcement learning framework for decision problems with risk objectives expressed through convex scoring functions. The framework can represent measures including variance, Expected Shortfall and entropic Value-at-Risk, as well as mean-risk utility. Its focus is on making these objectives workable over time, where optimizing risk can create time inconsistency: a policy preferred earlier may no longer be preferred at a later decision point.

To address this, the authors augment the state with an auxiliary variable and reformulate the problem as a two-state optimization. They describe a customized Actor-Critic algorithm, theoretical approximation guarantees that do not require a continuous Markov decision process, and an auxiliary-variable sampling procedure inspired by alternating minimization. The paper reports simulation experiments, including a statistical arbitrage application, as evidence of the method’s effectiveness. The available description gives no numerical results or detailed experimental setup, and the sampling method’s convergence is conditional on stated assumptions; practical performance beyond the tested simulations is therefore unclear.

Key ideas

  • Convex scoring functions can express several familiar risk measures within a reinforcement learning objective.
  • Augmenting the state with an auxiliary variable helps address time inconsistency in risk-sensitive decisions.
  • The proposed Actor-Critic method comes with approximation guarantees that do not require a continuous state process.
  • An alternating-minimization-inspired sampling method is reported to converge under certain conditions.
  • Simulation experiments include an application to statistical arbitrage trading.

Tags

Full text
# Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions


# Risk-sensitive Reinforcement Learning Based on Convex Scoring Functions









We propose a reinforcement learning (RL) framework under a broad class of risk objectives, characterized by convex scoring functions. This class covers many common risk measures, such as variance, Expected Shortfall, entropic Value-at-Risk, and mean-risk utility. To resolve the time-inconsistency issue, we consider an augmented state space and an auxiliary variable and recast the problem as a two-state optimization problem. We propose a customized Actor-Critic algorithm and establish some theoretical approximation guarantees. A key theoretical contribution is that our results do not require the Markov decision process to be continuous. Additionally, we propose an auxiliary variable sampling method inspired by the alternating minimization algorithm, which is convergent under certain conditions. We validate our approach in simulation experiments with a financial application in statistical arbitrage trading, demonstrating the effectiveness of the algorithm.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.