Skip to content
All library documents

Why Strong Classification Scores Can Still Lose After Trading Costs

Article Quant Q&A · Author: PyRsquared

Summary

The document describes a trading system that predicts buy or sell labels and sizes positions using predicted probabilities. Despite similar out-of-sample F1 scores in training and backtest periods, the strategy loses money after commissions and slippage. The author uses OHLCV features and labels observations with a triple-barrier method: returns are classified according to whether they hit volatility-scaled upper or lower barriers within a fixed horizon, with a time-out assigned a sell label.

The response recommends inspecting individual trades and asks whether the simulation is event-driven or a simple loop. It points out that a model restricted to buy or sell cannot choose to hold, potentially causing frequent trading and costs even when expected price moves are small. The central practical point is that classification quality does not establish net profitability: trading costs must influence decision-making. The short answer suggests making costs endogenous to the model but does not provide a specific method, and the reported F1 score alone cannot show whether predicted returns exceed costs.

Key ideas

  • A strong out-of-sample classification score does not guarantee that a strategy will be profitable after execution costs.
  • Trade-by-trade inspection can help identify whether simulated execution or frequent low-value trades drive losses.
  • A forced buy-or-sell decision removes the option to hold and may increase unnecessary turnover.
  • Triple-barrier labels use volatility-scaled price thresholds and a time limit to classify outcomes.
  • Trading costs need to affect trade selection or sizing, rather than be assessed only through classification metrics.

Tags

Full text
# Accurate model but execution in backtesting is losing money


# Accurate model but execution in backtesting is losing money












I have a binary classification model that predicts BUY (1) and SELL (-1) with an out of sample F1 score of 71% (precision is 65% and recall is 80%). The model's output is a probability of a BUY label occurring, which is then used in a bet sizing formula to bet cash (similar to the kelly criterion). The model was also trained taking commission and conservative slippage into account.

Now that I have the trained model, I am backtesting on out of sample data. However, it seems that the execution is losing money in the long term. The model is performing as expected - even on the backtest data, it still performs with an F1 score of 71%, but it seems commission and slippage are eroding gains.

Are there any tips / rules / literature one should follow when backtesting and simulating execution? I find it strange that the model performs so well on out of sample data yet still loses money - again, the model was trained with conservative estimates on slippage, higher than would ever occur in actual trading.

EDIT: As suggested in the comments, I will expand a bit more on what the model is doing. The data is simple OHLCV time-series with some feature engineering applied (e.g. CDF values for the distribution of returns, making the prices stationary through the backshift operator, normalizing data between 0 and 1). The method for labelling data is taken from Lopez's Advances in Financial Machine Learing, specifically the Triple Barrier Labelling method; first you calculate percentage returns from close price, then calculate a EWMA of standard deviation (std) of these returns - this is like an implied volatility. The upper barrier is a multiple of the EWMA of std of returns and similarly with the lower barrier, the 3rd vertical barrier is a fixed window of time later (say 60 minutes, or 5 days). At each time `i`, calculate the returns between `i` and `i + j`. If this return hits (is greater than or equal to) the upper barrier (upper EWMA of std of returns at `i`), it is labelled BUY (1) at time `i`. If the return hits the lower barrier (lower EWMA of std of returns at `i`), it is labelled SELL (-1) at time `i`. If the return has not hit either upper or lower barrier by the time the fixed window time has elapsed, it is labelled a SELL (-1) at time `i`.

The backtest data is distributed the same as the training data, as shown by a chi-square test and the fact that the model achieves similar F1 scores between test data and backtest data (both out of sample).

## Answer by Hamish Gibson (score 2)

https://quant.stackexchange.com/a/53470

Can I ask what type of backetest you are using? Is it an event-driven backtest or a simple for-loop backtest. Depending on whether you wrote your own or are using a library for it, try and analyse a sample of trades individually.

If you are only predicting a buy or sell, then you aren't really left with any wriggle room such as a `hold` option. So therefore each trade you are making means you are paying transaction fees even if the price movements don't warrant any significance. I am building a system too which has any profits being whittled away by fees. You have to find a way to somehow make the concept of fees become something endogenous to your ML model, because granted whilst 71% is impressive, trading fees are something exogenous to ML models.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.