Adapting Differential Sharpe Rewards to Downside Risk
Summary
The document asks how to adapt the differential Sharpe ratio reward used in a trading reinforcement learning system so that it instead reflects the Sortino ratio. The proposed motivation is that the traded assets are volatile and positively skewed, making downside-focused risk adjustment seem more suitable than penalizing total return variance. The author describes using the stepwise influence of each return on the cumulative Sharpe ratio and wants an analogous measure based only on downside risk.
No derivation, implementation, or empirical test is included; the text is a request for guidance. It therefore raises the objective-design problem rather than demonstrating a solution. Any adaptation would need to define downside deviation and its reference threshold consistently, then assess whether the resulting per-step reward behaves as intended during learning. The reported Sharpe and Sortino figures are specific to the author’s strategy and do not establish that Sortino optimization is generally preferable.
Key ideas
- Differential Sharpe rewards attribute each step’s return to a changing cumulative risk-adjusted measure.
- The author seeks a corresponding reward based on downside variability for Sortino optimization.
- The motivation is that positive skew can make total volatility a less fitting penalty than downside risk.
- The document poses the derivation question but does not provide a method or validation.
Tags
Full text
# Differential Sortino Ratio # Differential Sortino Ratio I'm attempting to optimize a reinforcement learning system to maximize risk adjusted returns. I have currently defined the reward as the differential Sharpe ratio at each step: the influence of the return at time t on the cumulative Sharpe ratio. This was defined in this paper [Reinforcement Learning for Trading, by John Moody and Matthew Saffell, NIPS, 1999], however I am interested in maximizing Sortino ratio instead. The assets I am trading are highly volatile but also heavily skewed towards positive returns; the Sharpe ratio of the strategy is only 2.5 but the Sortino is ~11. So, maximizing for Sharpe here doesn't make much sense. The derivation of this differential Sharpe ratio is: I would like to adapt this only take into account the variance of my downside risk. Any help or further reading would be great.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.