Using Exponential Averages to Approximate Sharpe and Sortino Ratios
Summary
The document examines differential Sharpe and downside deviation ratios (DSR and D3R), which use exponentially weighted estimates of returns and their squares to update risk-adjusted performance measures incrementally. This can avoid recalculating a full-history ratio at every time step, though the EMA approach gives more weight to recent observations than an online estimate over the entire history.
The author reports that extreme or unstable estimates came from insufficient return scaling, limited floating-point precision, and an adaptation rate that was too high. Lowering that rate and scaling simple returns produced more plausible estimates in their experience. They also found logarithmic returns problematic for their Sortino estimate because they treat gains and losses asymmetrically. The discussion is an implementation troubleshooting report, not a controlled comparison; it does not establish universal parameter choices or that the proposed adjustments suit every return series. It also raises startup instability and initialization as open practical questions.
Key ideas
- Differential Sharpe and downside deviation ratios can update performance estimates from exponentially weighted return moments.
- The adaptation rate controls responsiveness and can cause unstable estimates when set too high.
- Return scaling may help avoid precision problems in floating-point calculations.
- The author reports that logarithmic returns produced undesirable Sortino behavior in their application.
- Online full-history estimates avoid the recency weighting inherent in exponential moving averages.
Tags
Full text
# Approximating Sharpe and Sortino ratios from Exponential moving averages
# Approximating Sharpe and Sortino ratios from Exponential moving averages
So I've been studying the paper "Learning To Trade via Direct Reinforcement" Moody and Saffell (2001) which describes in detail how to use exponential moving estimates (EMAs) of returns at time t (`r_t`) to approximate both the Sharpe and Sortino ratios for a portfolio or security.
Note: in the paper he refers to the Sortino ratio as the "Downside Deviation Ratio" or DDR. I'm quite certain that mathematically speaking there is no difference between the DDR and the Sortino ratio.
So, the paper defines two values used to approximate either ratio, the Differential Sharpe Ratio (`dsr`) and the Differential Downside Deviation Ratio (`d3r`). These are calculations which both represent the influence of the trading return at time `t` (`r_t`) on the Sharpe and Sortino ratios. The EMAs used to calculate the DSR and D3R are based on an expansion around an adaption rate, `η`.
He then presents an equation by which I should be able to use the DSR or D3R at time `t` to recursively calculate a moving approximation of the current Sharpe or Sortino ratios at time `t` without having to perform a calculation over all t to get the exact result. This is very convenient in an environment with an infinite time horizon. Computationally, the data eventually would get too big to recalculate the full Sharpe or Sortino ratio at each timestep `t` if there are millions of timesteps.
$$S_t |_{\eta>0} \approx S_t|_{\eta=0} + \eta\frac{\partial S_t}{\partial \eta}|_{\eta=0} + O(\eta^2) = S_{t-1} + \eta\frac{\partial S_t}{\partial \eta}|_{\eta=0} + O(\eta^2)$$ $$D_t \equiv \frac{\partial S_t}{\partial \eta} = \frac{B_{t-1}\Delta A_t - \frac{1}{2}A_{t-1}\Delta B_t}{(B_{t-1} - A_{t-1}^2)^{3/2}}$$ $$A_t = A_{t-1} + \eta \Delta A_t = A_{t-1} + \eta (R_t - A_{t-1})$$ $$B_t = B_{t-1} + \eta \Delta B_t = B_{t-1} + \eta (R_t^2 - B_{t-1})$$
Above is the equation to use the DSR to calculate the Sharpe ratio at time `t`. To my mind, larger values of `η` might cause more fluctuation in the approximation as it would put more "weight" on the most recent values for `r_t`, but in general the Sharpe and Sortino ratios should still give logical results. What I instead find is that adjusting `η` wildly changes the approximation, giving totally illogical values for the Sharpe (or Sortino) Ratios.
Similarly, the following equations are for the D3R and approximating the DDR (a.k.a Sortino ratio) from it:
$$DDR_t \approx DDR_{t-1} + \eta \frac{\partial DDR_t}{\partial \eta}|_{\eta=0} + O(\eta^2)$$ $$D_t \equiv \frac{\partial DDR_t}{\partial \eta} = \\ \begin{cases} \frac{R_t - \frac{1}{2}A_{t-1}}{DD_{t-1}} & \text{if $R_t > 0$} \\ \frac{DD_{t-1}^2 \cdot (R_t - \frac{1}{2}A_{t-1}) - \frac{1}{2}A_{t-1}R_t^2}{DD_{t-1}^3} & \text{if $R_t \leq 0$} \end{cases}$$ $$A_t = A_{t-1} + \eta (R_t - A_{t-1})$$ $$DD_t^2 = DD_{t-1}^2 + \eta (\min\{R_t, 0\}^2 - DD_{t-1}^2)$$
I wonder if I'm misinterpreting these calculations? Here's my Python code for both risk approximations where `η` is `self.ram_adaption`:
```
def _tiny():
return np.finfo('float64').eps
def calculate_d3r(rt, last_vt, last_ddt):
x = (rt - 0.5*last_vt) / (last_ddt + _tiny())
y = ((last_ddt**2)*(rt - 0.5*last_vt) - 0.5*last_vt*(rt**2)) / (last_ddt**3 + _tiny())
return (x,y)
def calculate_dsr(rt, last_vt, last_wt):
delta_vt = rt - last_vt
delta_wt = rt**2 - last_wt
return (last_wt * delta_vt - 0.5 * last_vt * delta_wt) / ((last_wt - last_vt**2)**(3/2) + _tiny())
rt = np.log(rt)
dsr = calculate_dsr(rt, self.last_vt, self.last_wt)
d3r_cond1, d3r_cond2 = calculate_d3r(rt, self.last_vt, self.last_ddt)
d3r = d3r_cond1 if (rt > 0) else d3r_cond2
self.last_vt += self.ram_adaption * (rt - self.last_vt)
self.last_wt += self.ram_adaption * (rt**2 - self.last_wt)
self.last_dt2 += self.ram_adaption * (np.minimum(rt, 0)**2 - self.last_dt2)
self.last_ddt = math.sqrt(self.last_dt2)
self.last_sr += self.ram_adaption * dsr
self.last_ddr += self.ram_adaption * d3r
```
Note that my `rt` has a value that oscillates around `1.0` where values `>1` mean profits and `<1` mean losses (while a perfect `1.0` means no change). I first make `rt` into logarithmic returns by taking the natural log. `_tiny()` is just a very small value (something like `2e-16`) to avoid division by zero.
My problem(s) are:
- I would expect the approximated Sharpe and Sortino ratios to fall in the range 0.0 to 3.0 (give or take) and instead I get a monotonically decreasing Sortino ratio, and a Sharpe ratio that can explode to huge values (over 100) depending on my adaption rate `η`. The adaption rate `η` should affect noise in the approximation but not make it explode like that.
- The D3R is (on average) negative more than it is positive, and ends up approximating a sortino ratio that falls in a near-linear manner, which if left to iterate for long enough can reach totally nonsensical values like -1000.
- There are occasionally very large jumps in the approximation which I feel could only be explained by some error in my calculations. The approximated Sharpe and Sortino ratios should have a somewhat noisy but steady evolution without massive jumps such as those seen in my graphs.
Finally, if someone knows where I could find other existing code implementations wherein the DSR or D3R is used to approximate the Sharpe/Sortino ratios it would be much appreciated. I was able to find this page from AchillesJJ but it doesn't really follow the equations put forth by Moody, as he is recalculating the full average for all previous timesteps to arrive at the DSR for each timestep `t`. The core idea is being able to avoid doing that by using the Exponential Moving Averages.
## Answer by Alex Pilafian (score 2, accepted)
https://quant.stackexchange.com/a/57654
For anybody still following this:
I figured out that the equations and my code work fine; the problem was that I had to scale the returns before doing the risk calculations to avoid float32 precision data loss, and also just that my value for `η` was far too high. Lowering my `η` value to `<= 0.0001` produces totally logical sharpe and sortino approximations. As a sidenote, this also allows my neural network to learn directly from the marginal sharpe and sortino calculations, which is great.
As well, using logarithmic returns was problematic for the sortino approximation, so I effectively changed it to `rt = (rt - 1) * scaling_factor` which makes the sortino approximation not tend towards negative values anymore.
Logarithmic returns would have worked fine if my only goal was to use the DSR/D3R as a loss calculation in my neural network, but to get good sortino approximations it doesn't work as it sharply emphasizes negative returns and smooths positive returns.
## Answer by babelproofreader (score 1)
https://quant.stackexchange.com/a/57581
If your concern is about computational efficiency in calculating Sharpe/Sortino over large and increasing amounts of data, you can use incremental/online methods to calculate means, standard deviations etc. over the whole data set. Then just use the latest, online calculated value for the Sharpe/Sortino of the whole data set. This will avoid the problem of older data having less weight than newer data, which is implicit when using EMAs.
My answer on the Data Science SE at https://datascience.stackexchange.com/questions/77470/how-to-perform-a-running-moving-standardization-for-feature-scaling-of-a-growi/77476#77476 gives more detail and a link.
## Answer by orie (score 0)
https://quant.stackexchange.com/a/60791
This has been really, really useful, thank you. I've applied this to an RL algorithm (just the DSR metric) and I have a few things to ask if this thread is still active.
- What do you do about the first steps? it seems like the values are unstable at the beginning of the sequence.
- Also, at what values would you initiate the moving averages?
- I've also experienced sudden drop during training
Why do you think that is?
Here's your code, just changed the naming and put it into a class, I hope I did it right
> class DifferentialSharpeRatio: def init(self, eta=1e-4): self.eta = eta self.last_A = 0 self.last_B = 0 `def _differential_sharpe_ratio(self, rt, eps=np.finfo('float64').eps): delta_A = rt - self.last_A delta_B = rt**2 - self.last_B top = self.last_B * delta_A - 0.5 * self.last_A * delta_B bottom = (self.last_B - self.last_A**2)**(3 / 2) + eps return (top / bottom)[0] def get_reward(self, portfolio): net_worths = [nw['net_worth'] for nw in portfolio.performance.values()][-2:] rt = pd.Series(net_worths).pct_change().dropna().add(1).apply(np.log).values dsr = self._differential_sharpe_ratio(rt) self.last_A += self.eta * (rt - self.last_A) self.last_B += self.eta * (rt**2 - self.last_B) return dsr `
```
def _differential_sharpe_ratio(self, rt, eps=np.finfo('float64').eps):
delta_A = rt - self.last_A
delta_B = rt**2 - self.last_B
top = self.last_B * delta_A - 0.5 * self.last_A * delta_B
bottom = (self.last_B - self.last_A**2)**(3 / 2) + eps
return (top / bottom)[0]
def get_reward(self, portfolio):
net_worths = [nw['net_worth'] for nw in portfolio.performance.values()][-2:]
rt = pd.Series(net_worths).pct_change().dropna().add(1).apply(np.log).values
dsr = self._differential_sharpe_ratio(rt)
self.last_A += self.eta * (rt - self.last_A)
self.last_B += self.eta * (rt**2 - self.last_B)
return dsr
```Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.