Skip to content
All library documents

Interpreting Probability of Backtest Overfitting Metrics

Article Quant Q&A · Author: mr.T

Summary

The document introduces the Probability of Backtest Overfitting (PBO) algorithm and asks how to interpret its reported metrics: p_bo, slope, adjusted R-squared, and p_loss. It includes one set of outputs from an example and lists the author’s proposed directions for several metrics, then provides a second example using synthetic profitability data divided into partitions and evaluated with an Omega-based function.

The second output is presented as a reproducible calculation, but the document does not supply an interpretation or an ideal target profile for a strategy without overfitting. It therefore illustrates how the package is called and what it returns, rather than explaining the metrics’ definitions or validating thresholds. The synthetic series and the stated function settings are limited examples; conclusions about a real strategy would require understanding the method’s assumptions, data partitioning, and the meaning of each reported statistic.

Key ideas

  • PBO is presented as a method for assessing whether a trading strategy’s backtest may be overfit.
  • The reported metrics include p_bo, slope, adjusted R-squared, and p_loss.
  • The document shows outputs from an example and a synthetic profitability series evaluated across partitions.
  • It asks for ideal metric values but does not provide definitions, interpretation, or recommended thresholds.

Tags

Full text
# PBO algorithm "The Probability of Backtest Overfitting" paper


# PBO algorithm "The Probability of Backtest Overfitting" paper












In this article by Lopez de Prado et al., an algorithm was proposed for assessing the overfit of a trading strategy:

The Probability of Backtest Overfitting

There is also a package for R: pbo: Probability of Backtest Overfitting

The following is the result of applying the function to an example:

```
##       p_bo      slope       ar^2     p_loss 
##  1.0000000 -0.0030456  0.9700000  1.0000000
```

I would like to know if I am interpreting the calculation result correctly.

`p_bo` - should go to zero

`slope` - should tend to 1

`ar^2` - should tend to 1

`p_loss` - should go to zero

############ UPD ############

Here is a reproducible code example.

This is what I do to evaluate the profitability of my trading strategy. I would like to know how to interpret the `PBO_metrics` result.

```
library(pbo)
library(PerformanceAnalytics) 

# profitability of a trading strategy
p <- cumsum(rnorm(5001)) + seq(0,200,length.out=5001)
plot(p,t="l",main="profitability of a trading strategy")

PBO_metrics <- diff(p) |> matrix(ncol = 20) |> as.data.frame() |> 
               pbo(s=8,f=Omega, threshold=1) |> 
               summary()
```

..

```
PBO_metrics
   p_bo   slope    ar^2  p_loss 
 0.3000  1.6049 -0.0150  0.1430
```

In other words, what values should an ideal non-overfitted trading strategy?

## Answer by Yagiz Temizel (score 0)

https://quant.stackexchange.com/a/85887

Yeah, your interpretation is right on all four.

pbo should go to 0 - it's literally the fraction of combinatorial splits where the strategy that looked best in-sample ends up at or below the out-of-sample median (that's the logit <= 0 case). Close to 0 means the in-sample winner usually keeps winning out-of-sample; close to 1 means picking the "best" in-sample strategy is basically a coin flip or worse.

slope and ar2 come from the same place - a linear fit between your in-sample and out-of-sample performance pairs across all the combinations. slope near 1 means performance barely degrades going from IS to OOS (good). ar2 near 1 means that relationship is actually tight and predictable, not just a loose scatter - so a strategy can have a decent slope but low ar2, which would mean the degradation itself is kind of all over the place and hard to trust as a clean signal. You want both high.

p_loss going to 0 just means that across your OOS splits, the strategy rarely dips below whatever loss threshold you set it against - so a smaller p_loss is better, same direction as pbo.

One thing worth flagging with your example code: you're generating p as a pure random walk with upward drift (cumsum(rnorm) + linear trend), so basically every slice of it will look "good" in a similar way regardless of split - that's not really testing for overfitting, it's testing whether a trending random walk looks consistent with itself, which it mostly will. To actually see pbo do something interesting you'd want several genuinely different return series (e.g. simulate a few where the apparent edge is concentrated in a handful of lucky periods) and feed those in as your strategy columns - that's when you'll see pbo climb toward 1 for the ones that are overfit.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.