Handling Trade Clusters Without Look-Ahead Bias in Backtests
Summary
The post considers how to compare profit factors across forex strategy runs when trades arrive in bursts. The author proposes treating nearby orders as one cluster and weighting each trade according to the number of subsequent orders within a short window. A reply explains that this uses future information unavailable when the first trade is placed, creating look-ahead bias and making the backtest unrealistic.
The suggested remedy is to model trade sizing using only information available at entry. If a strategy cannot forecast clustering, one alternative is to allocate capital incrementally across successive trades, reducing the amount assigned to each later entry. Another reply suggests aggregating price data into candlestick summaries and using high, low, and close values to produce distinct profit-and-loss estimates, including a pessimistic scenario for a long trade. These are possible evaluation approaches, not a validated universal clustering method; results depend on the strategy’s execution and capital constraints.
Key ideas
- A weighting rule based on future trades in a cluster introduces look-ahead bias.
- Backtests should reflect the information and capital available when each trade is entered.
- Incremental capital allocation can model the diminishing capacity to take additional clustered trades.
- Candlestick high, low, and close values can support multiple profit-and-loss estimates.
Tags
Full text
# How to "uncluster" a set of financial data? # How to "uncluster" a set of financial data? I am attempting to evaluate and compare the profit factor of different "test runs" of a FOREX trading strategy. My problem is that, despite an average time between orders of 2hr+, some of these runs can have 20+ orders in a row, every 5 minutes, in the same direction. I need some way to normalize these clusters that occur without just throwing the data out. I want to treat the cluster as 1 data point by averaging the gain/loss of each trade within the cluster. I was thinking of doing it with a moving time window in the following manner: ``` For each order: Weight = 1/(N+1) N = count of consecutive orders within 30 minutes of the current order. ``` But I am not sure if that is correct. ## Answer by Tal Fishman (score 5) https://quant.stackexchange.com/a/1953 That is definitely not correct. Your test results will suffer from look-ahead bias. At the point in time that you will be entering your order, you will not know yet how many orders will come in the next 30 minutes, thus a proper backtest should not use that information to size the trade. There are many factors to consider when running a proper backtest (see wikipedia), but the key factor behind all of them is to replicate as closely as possible what is realizable in live trading. In your case, that means that you will have to predict how many trades your strategy is about to enter and size all subsequent trades appropriately. Once you reach your "limit," at which point in real trading you would have run out of capital, your backtest is effectively prevented from entering new trades. If your strategy isn't able to predict whether trades are about to cluster, you could try an incremental system, whereby the first trade is assigned a proportion $\delta$ of total capital, the second gets $\delta(1-\delta)$, etc. ## Answer by icequations (score 2) https://quant.stackexchange.com/a/2225 you could perhaps cluster the information in a candle stick manner and bin the price data into high, low and closing, instead of throwing some data out and keeping only its closing price. Using the high, low and closing prices you are able to make 3 separate estimates with regards to Profit and lost. Obviously with a long position, profit calculated by entering at a high price while using a low price for its exit would be the pessimistic estimate of profit.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.