Skip to content
All library documents

Policy Gradient Reinforcement Learning with Dynamic Convex Risk Measures

Article arXiv papers · Author: Anthony Coache et al.

Summary

The document presents a model-free reinforcement learning method for sequential optimization when outcomes are evaluated using dynamic convex risk measures. It uses a time-consistent dynamic programming principle to value policies, then derives policy-gradient updates for finding improved policies. An actor-critic approach with neural networks is proposed to optimize those policies.

The method is demonstrated on three tasks: statistical arbitrage trading, financial hedging, and obstacle-avoidance control for a robot. This establishes that the framework is applied across both financial and nonfinancial settings, but the description supplies no performance figures, comparison methods, market assumptions, or implementation details. As a result, it conveys the optimization structure and application scope, while leaving evidence about trading effectiveness and practical limitations unspecified.

Key ideas

  • Dynamic convex risk measures provide a way to evaluate sequences of uncertain outcomes.
  • Time-consistent dynamic programming is used to value candidate policies.
  • Policy-gradient rules and a neural-network actor-critic method support policy optimization.
  • Demonstrations include statistical arbitrage, financial hedging, and robot obstacle avoidance.
  • The available description gives no comparative results or market implementation assumptions.

Tags

Full text
# Reinforcement Learning with Dynamic Convex Risk Measures


# Reinforcement Learning with Dynamic Convex Risk Measures









We develop an approach for solving time-consistent risk-sensitive stochastic optimization problems using model-free reinforcement learning (RL). Specifically, we assume agents assess the risk of a sequence of random variables using dynamic convex risk measures. We employ a time-consistent dynamic programming principle to determine the value of a particular policy, and develop policy gradient update rules that aid in obtaining optimal policies. We further develop an actor-critic style algorithm using neural networks to optimize over policies. Finally, we demonstrate the performance and flexibility of our approach by applying it to three optimization problems: statistical arbitrage trading strategies, financial hedging, and obstacle avoidance robot control.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.