Why Backtest Sharpe Ratios Need Haircuts and Realistic Execution
Summary
The discussion asks whether backtested Sharpe ratios should routinely be cut in half. The answers describe that figure as a rough and discretionary haircut, not a universal statistical rule. Strategies discovered through broad data searches may be less persistent than strategies grounded in an advance rationale, because selection can make historical performance look unusually strong. For active strategies, the information ratio may better isolate returns beyond factor exposures, but it too can be inflated by research choices.
A separate source of optimism is execution: historical tests may misstate fill prices, fill probabilities, liquidity costs, and the market impact of orders. Conservative assumptions about crossing the spread or placing limits can reduce estimated performance, while unfilled orders and opportunity costs also matter. Borrow costs, forced short covering, unexpected expenses, and competition can further erode returns. The document offers practical cautions rather than a single validation recipe; it does not establish that any particular haircut makes a backtest reliable.
Key ideas
- A 50% Sharpe haircut is presented as a discretionary heuristic, not a universal rule.
- Strategies found through data mining may be less persistent than those based on prior rationale.
- The information ratio can focus evaluation on alpha after accounting for factor exposures, but it can still be overstated.
- Backtests should model fills, liquidity, market impact, unfilled orders, and shorting costs.
- Realized strategy performance may decline as competitors discover similar signals.
Tags
Full text
# Why discount results by 50%? # Why discount results by 50%? I am reading Harvey & Liu (2015) article, which says the following: > A common practice in evaluating backtests of trading strategies is to discount the reported Sharpe ratios by 50%. There are good economic and statistical reasons for reducing the Sharpe ratios. The discount is a result of data mining. This mining may manifest itself by academic researchers searching for asset pricing factors to explain the behavior of equity returns or by researchers at firms that specialize in quantitative equity strategies trying to develop profitable systematic strategies. If I understand it correctly, if Strategy A is showing Sharpe Ration of 10, I would adjust it to 5. I am not sure that I really understand why it is a common practice to discount. Probably one wants to show a more pessimistic results, given that the results are backed by historical data? Also, in the paper authors introduce some t-test approaches to test validity of the backtest results. Are there other common methodologies to verify that backtest results are trustworthy? ## Answer by kurtosis (score 1, accepted) https://quant.stackexchange.com/a/80046 Harvey and Liu distinguish between trading strategies with theoretical backing versus those which are suggested merely by empirical analysis. The 50% discount of Sharpe ratio is completely arbitrary. Furthermore, I would argue the real figure could vary a lot from even that discount. First, if you have a quantitative active strategy, a Sharpe ratio is a poor metric: you should be contributing alpha, not beta. That means you should not be looking at the Sharpe ratio -- which uses absolute returns beyond a risk-free rate divided by the return volatility. Instead, you should look at the information ratio -- which has the alpha (mean abnormal returns, after subtracting the strategy's exposure to some factors) divided by the standard deviation of that alpha. Is that information ratio still trustworthy? No, that also should be discounted. Second, how much to discount these metrics is imprecise: it's a hack to accommodate the reality than many backtests are not accurate reflections of likely fills were you to really trade. Estimating fill probabilities and prices is tough without actually trading the strategy. I've seen strategies which make money in backtests but lose money consistently and quickly once put into production. The problem is liquidity: what will it cost you to trade when you actually get your information and decide to trade? A backtest needs to know a lot about a market to even estimate if you might expect to get your order filled. Worse still: your order can alter the distribution of people placing orders on the other side of yours. (In the most extreme example: your order might scare away orders on the other side leaving you unfilled.) Most traders I know with trading strategies either assume (1) they have to cross the spread and be less than the size on the far side when they place their order, or (2) that they would place a limit order and the market would have to trade at a price better than their limit price for a size greater than their order size -- and unfilled orders are tracked to measure opportunity cost. Third: Lets say you have these (hopefully) conservative estimates of fill prices and probabilities; and, that despite these your strategy would still be profitable. Those conservative estimates can often yield a substantial reduction in Sharpe and information ratios. You might still discount those ratios further because you do not know how your order will change the market against your favor -- and you might incur unforeseen expenses. (For example: often stocks you want to short are also in shorting demand from others; fees might be higher for borrow and you might even get kicked out of your short prematurely.) Finally, as you trade you should expect your alpha to decay over time as others sniff out your strategy (or strategies like it) and as markets become more efficient. Building a backtest which is even somewhat realistic is one of the more difficult tasks in finance. Using that to then assess potential trading strategies is even harder. Hence why you see so many conservative assumptions: the task is hard and reality can often be far harsher. If you want a good longer read on this topic, I would recommend Brian Peterson's Developing & Backtesting Systematic Trading Strategies. ## Answer by KaiSqDist (score 2) https://quant.stackexchange.com/a/80031 Previously, I thought this was an excellent question and very practical/applicable to performance analytics. After some related readings, I feel that I am (somewhat) ready to provide a (tiny) insight to this discussion. In Systematic Trading by Rob Carver, the author talks about two kinds of trading rules - data first/data mining and ideas first. For data mining, it is about looking at data (even big data) and analyzing them to discover profitable patterns to trade on. These patterns are are not backed by ideas or logic, and are discovered first then backed up by logic to explain these profitable patterns ex-post. Therefore, it is in my opinion that it is more likely that these profitable patterns from data mining are less persistent than an ideas first technique. Thus, it makes sense for practitioners to discount the Sharpe ratios by 50% or another significant amount. I am not entirely sure WHY the amount chosen is 50%, but there must be empirical evidence to suggest this number but it seems like the percentage chosen is more of a discretionary nature. UPDATE I would like to add that this 50% should be explained in page 90 of Rob Carver's book under Table 14, where the author explains different discount amounts (to be applied on back-tested Sharpe ratios) based on the specific fitting technique used in back-tests.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.