Evaluating Backtests with Returns, Capital, and Trading Constraints
Summary
The discussion compares evaluating a strategy through trade or portfolio returns with simulating a starting capital amount and measuring ending equity or annualized return. It argues that the appropriate view depends on the strategy and how closely the backtest must represent deployable trading. A capital-based simulation can capture contract sizes, position limits, minimum ticks, and other constraints that return-only calculations may not represent reliably.
The answers also recommend looking beyond cumulative return. Relevant measures include drawdowns, trade counts, holding periods, index correlation, Sharpe ratio, commissions, and turnover. Time-weighted and dollar-weighted returns can each be informative, while leverage and margin calls make drawdown analysis essential. The discussion notes that portfolio starting conditions can affect what counts as good performance, and that simulation granularity, trading costs, and adequate capital matter. These are general considerations rather than a single prescribed evaluation framework; the right metrics depend on the strategy, instrument, and intended use.
Key ideas
- Return-based results are convenient for comparing trades, while capital-based simulations can represent deployability constraints.
- Futures contract sizes, position limits, and minimum tick movements can make capital assumptions important.
- Backtest evaluation should consider drawdowns, turnover, costs, trade behavior, and risk-adjusted performance alongside returns.
- Time-weighted and dollar-weighted returns answer different questions and can both be useful.
- Simulation detail and starting portfolio conditions affect how closely results reflect live trading.
Tags
Full text
# How to properly evaluate backtest returns? # How to properly evaluate backtest returns? Do you evaluate a strategy in a backtest based on the cumulative returns generated by the strategy (i.e. looking at the cumulative returns of the trades that occur) or do you start with a certain dollar amount and look at the cash at end to compute the annualized return. The reason I ask is because things like position sizing and adding position to a certain trade are much easier to do when starting out with a dollar amount in a portfolio. I tend to look at a strategy as a collection of trades with particular entry and exit points. If using multiple entry & exit points are used for each signal (adding positions or taking off positions for each signal) it is much easier to calculate returns per trade if looking at a dollar amount that was invested due to each signal. I have seen both approaches being used and personally prefer using returns rather than starting with a dollar amount. Just curious to see how people approach this and why. ## Answer by hroptatyr (score 7, accepted) https://quant.stackexchange.com/a/2519 I'd say it depends on how close you want to be to reality and what the strategy entails. For instance one scenario when actual currency makes sense is when you want to take contract sizes and position limits into account, for instance agricultural futures contracts nearly always impose a position limit for one party in one or all contracts. If your strategy on the other hand wants to invest, say, all returns in one or more such contracts the strategy will either fail, well, not in the backtest but in the real world. Representing hard limits like contracts sizes, position limits, minimum tick movement, etc. in terms of returns is impossible or at best questionable. I'm sure this scenario doesn't apply to you as you can obviously go with either approach, I just want to point this out to other people that may have a similar question in a different context. ## Answer by Steve Severance (score 6) https://quant.stackexchange.com/a/2518 We don't give strategies a dollar amount during backtesting, rather backtesting shows how much capital would be required to successfully deploy a strategy. We also don't look just at returns but many different metrics including but not limited to, max drawdown, variance of draw downs, number of trades, holding period, correlation (or lack of) with various indexes, Sharpe ratio, commissions paid, share turnover, and numerous other measures of variance. I view position sizing as a somewhat separate problem that we currently don't completely tackle at backtesting time. There may be max position sizes that a strategy can use for a given opportunity but often this is somewhat dependent on security being traded. ## Answer by Patrick Burns (score 6) https://quant.stackexchange.com/a/2521 http://www.portfolioprobe.com/2010/11/05/backtesting-almost-wordless/ shows an example of how the results from a backtest can be deceiving. This would be true with either returns or value. The main issue is that the portfolio you start with can have an impact on what "good" means. ## Answer by ProbablePattern (score 4) https://quant.stackexchange.com/a/2517 I think your question may be getting at the difference between time weighted returns and dollar weighted returns. The only advantage I see to using a dollar amount is to simplify drawdown calculations. Both dollar weighted and time weighted returns are necessary to evaluate a strategy particularly when benchmarking against a competing strategy. Ignoring drawdowns by looking at a starting dollar amount and ending dollar amount will become problematic when you want to leverage the strategy and margin calls may occur. ## Answer by Jon Grah (score 1) https://quant.stackexchange.com/a/32618 @hroptatyr, @Steve, @Tal Fishman, and others here have alluded to the main problem of evaluating backtests. First we have to define what constitutes a backtest, which is simply an ability to simulate market activity and how specific instructions (algo) would perform in that environment. The simulation is supposed to best reconstruct how the market would have behaved at the time. Tick level simulation is the most precise way to do this, but ms, second, and minute data is increasingly less precise, but faster to process. A strategy where intraday timing may be able to get away with minute backtesting, but more than likely you would have to at least use seconds interval as minimum to properly simulate what would have happened. Then you need to know what the algo system is supposed to do. In other words, What category of speculation is it? (e.g. news trading, arbitrage, value trading, dealer, etc) did the algo actually perform (execute its functions) as advertised. A lot of investors who are not actually owners of the algo (black box) or have insight as to how the inputs work (grey box) will only have the usual results from the backtest (equity, total trades, sharpe ratios, etc). There is a structural conflict of interest there; you are just hoping that the live results will be similar without knowing why you were profitable. Easy to succumb to sample selection bias without covering the basics. Then you can determine the adequate starting capital and see if the returns net of estimated trading costs make it worthwhile.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.