Skip to content
All library documents

Cross-Asset Sampling and Price Versus Return Signals in Backtests

Article Quant Q&A · Author: monraf

Summary

The document discusses two design choices in strategy backtesting across a large, liquidity-filtered stock universe: selecting a random subset of assets for development and validation, and using prices or returns to generate signals. The response says random asset sampling can reduce the temptation to tune parameters to particular securities and lower computation demands. It cautions that this approach may be less representative if the intended live strategy will trade the full, specific universe; sample size and parameter complexity matter.

On signals, the response favors prices unless a strategy specifically depends on returns. Its reasoning is that trades and accounting occur at prices, while returns depend on the return definition and measurement interval. This is presented as a practical opinion, not a universal rule or empirical comparison. The document does not prescribe a sampling procedure, discuss repeated-sampling uncertainty, or address other backtest risks such as temporal leakage, transaction costs, or survivorship bias.

Key ideas

  • Randomly sampling assets can reduce tuning to a fixed set of securities and reduce computational cost.
  • A sampled-asset design may not represent a strategy intended for the entire target universe.
  • The appropriate sample size depends partly on the universe, parameter search, and live-trading objective.
  • Price signals correspond directly to transaction prices, while return signals depend on their definition and horizon.
  • Returns remain appropriate when the strategy explicitly uses return calculations.

Tags

Full text
# Backtesting Strategies - Sampling and returns types


# Backtesting Strategies - Sampling and returns types












I am back testing a strategy in R and I have some questions about testing design. I have a universe of around 500 stocks to test filtered based on liquidity. To test the trading strategy I have implemented a simple in-sample/out-of-sample testing scheme. Before running a full test on the whole universe of stocks I have been running the tests over a random sample of the stocks while designing a generic testing framework.

I have been dividing the universe into two parts, I have been optimizing over a random sample of different stocks then using the optimized parameters from the training set to test over a random selection of different stocks from the test set. Are there any problems with doing this rather than testing over the total universe?

My other question is to do with using prices vs returns (vs log returns). I am a little confused as to the reasoning to use returns rather than price to find signals. Would you use returns series because you are more able to compare returns series more easily then price series?

## Answer by Jacob Amos (score 0, accepted)

https://quant.stackexchange.com/a/25235

I think the random sampling has some definite advantages. First and foremost, it reduces the risk that you will be "overfitting" by optimizing your strategy's parameters to a certain sample of assets. Second, with such a large asset universe, it would obviously take tons of time and computational resources to test a strategy over historical data of any meaningful length, so you kind of avoid "punishing" yourself for using a long historical timeframe to test on.

That being said, if your end aim is to live trade the strategy on the those 500 assets specifically, maybe you lose something by only optimizing over a random sample of those assets. There's some critical thinking involved here as to how big the random sample should be, how large the parameter space, and so on. So at least part of the answer will depend on your end objectives.

As for using prices vs. signals, IMHO using returns can be really dangerous. Why? When you trade, transactions are done at a given price, not a given return (at least as far as the accounting is concerned). The other thing worth keeping in mind is that a return is dependent on:

(a) What kind of return (simple, geometric, log, etc.)

(b) When the calculation starts (1-year return, 1-day return, etc.)

Unless the strategy is contingent upon the calculation of returns, I would prefer the use of price data instead. (Since computational efficiency appears to be of concern to you given your large dataset, this also has the added advantage of requiring marginally fewer computations since you don't have to calculate returns in order to get your signals.)

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.