Skip to content
All library documents

Robust Risk-Aware Reinforcement Learning for Portfolio Strategies

Article arXiv papers · Author: Sebastian Jaimungal et al.

Summary

This work develops a reinforcement learning approach for optimizing policies under risk-aware performance criteria. It evaluates policies using rank-dependent expected utility, which lets an agent express different preferences over gains and downside outcomes. To address model uncertainty, the policy is assessed under the worst-case return distribution within a Wasserstein distance ball around the modeled distribution. This creates a nested decision problem: the agent selects a policy, then an adversary selects a nearby distribution that degrades its performance.

The authors derive policy gradient formulas for both the policy selection and adversarial distribution problems. They demonstrate the method on robust portfolio allocation, benchmark optimization, and statistical arbitrage. These applications show the framework’s intended scope, but the document gives no comparative performance figures, market data details, or practical implementation guidance. Its usefulness therefore depends on further validation of the model, uncertainty set, and risk preferences in a specific trading setting.

Key ideas

  • Rank-dependent expected utility allows policies to encode different trade-offs between gains and downside risk.
  • The framework evaluates each policy against a worst-case distribution within a Wasserstein neighborhood.
  • An adversarial inner problem seeks a distribution that reduces the chosen policy’s performance.
  • The authors derive policy gradients for both policy optimization and the adversarial problem.
  • Demonstrations cover portfolio allocation, benchmark optimization, and statistical arbitrage.

Tags

Full text
# Robust Risk-Aware Reinforcement Learning


# Robust Risk-Aware Reinforcement Learning









We present a reinforcement learning (RL) approach for robust optimisation of risk-aware performance criteria. To allow agents to express a wide variety of risk-reward profiles, we assess the value of a policy using rank dependent expected utility (RDEU). RDEU allows the agent to seek gains, while simultaneously protecting themselves against downside risk. To robustify optimal policies against model uncertainty, we assess a policy not by its distribution, but rather, by the worst possible distribution that lies within a Wasserstein ball around it. Thus, our problem formulation may be viewed as an actor/agent choosing a policy (the outer problem), and the adversary then acting to worsen the performance of that strategy (the inner problem). We develop explicit policy gradient formulae for the inner and outer problems, and show its efficacy on three prototypical financial problems: robust portfolio allocation, optimising a benchmark, and statistical arbitrage.

Shown in full with attribution under the source's licence. Licence: abstract CC0

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.