Skip to content
All library documents

Time-Ordered Validation for Financial Time-Series Models

Article Quant Q&A · Author: Michael Teguh Laksana

Summary

The discussion considers how to split tick data when tuning a fast trading strategy with machine learning. The main recommendation is to preserve chronological order: train and validate on earlier observations, then reserve a later period for testing. Randomly selecting short windows can produce a sample that does not represent the broader history, while combining random windows with chronological partitions is described as difficult to manage without introducing bias.

For model selection, the replies suggest time-series cross-validation, particularly rolling windows, to evaluate performance across different periods and limit overfitting. They also note that market regimes and seasonality can affect results: the training and test periods may differ, and recent data may sometimes be more relevant because of changed market conditions or technology. These are general suggestions rather than a tested comparison. The discussion does not prescribe a split ratio, account for label overlap or execution effects, or establish how much history is appropriate; choices depend on data quality and the model being used.

Key ideas

  • Chronological splits preserve the temporal order needed to evaluate a forecasting strategy.
  • Randomly sampled windows may fail to represent the full history and can create misleading evaluation results.
  • Rolling-window cross-validation can assess performance across multiple time periods.
  • Seasonality and changing market regimes can make training and test periods behave differently.
  • The suitable history and validation design depend on the data and the model.

Tags

Full text
# Train-test split configuration on timeseries data for machine learning optimization


# Train-test split configuration on timeseries data for machine learning optimization












I have a strategy (running in the seconds scale) which parameters I would like to optimize. The thing is I'm relatively new to financial machine learning and I'm not quite sure how to split the data for fitting the parameters in a general sense.

- What would be the general best practice for splitting the data for the given strategy's timescale? For example, if I have a tick data from 2009 to 2023, should I split the data 80:20 by:

a) time, say, around 2009-2020 as train/validation, 2021-2023 as test.

b) randomly sample a contiguous window (say 10 minutes) of data such that the total time split between the train and test are 80:20.

c) Combine a and b: Split between 2009-2020 as train/validation, 2021-2023 as test and do random sampling like in b when training and testing for each set of data.

d) any other recommendation?

- Should I be worried about seasonality in the data? For example if I go with option a) and the market is mainly bullish in the train data period, would it bias the model toward a bullish market and cause the model to perform worse on a bearish market? I assume option b) would solve this issue?

Thank you in advance

## Answer by FinEsta (score 0, accepted)

https://quant.stackexchange.com/a/80139

I would recommend using Option A. The reason is because Option B will introduce biases if the sampled windows are not representative of the broader time period and is thus not suitable for long time periods and Option C is difficult to manage and split to avoid bias since you'll have to check the split data to ensure representation across periods.

I would recommend doing Cross-Validation though; consider using cross-validation techniques tailored to time-series data, such as rolling-window cross-validation. This can help assess model performance across different time periods and reduce overfitting.

Specifically for Seasonality, if your data exhibits clear seasonality, ensure that these are reflected in both the training and test sets. This might involve adjusting the time-based split or ensuring that different regimes are represented in both sets.

## Answer by Andreu Boix Torres (score 0)

https://quant.stackexchange.com/a/80138

- That's an interesting question. Personally, I would recommend option A to avoid introducing potential biases to the model.

- The seasonality problem is indeed a concern. While there's no perfect solution, some strategies account for this by focusing on more recent data. For example, they might use data from 2015 to 2023 based on the assumption that the market conditions in this recent period are more relevant than those from 2009 to 2023, based on maybe relevance of historical event (such as covid) or technological advancements (such as AI, for instance).

Hope that helps, good luck with your strategy!

## Answer by Anuj (score 0)

https://quant.stackexchange.com/a/80140

- Option A is the best option to choose, however, it purely depends on the data quality and the machine learning or classical time series model you are trying to implement.

- Another method you may implement is by using k-folds and hypertuning with grid-search so that your model will learn the variance and pattern of the data. Then use the Train, Validate and Forecast method.

- Furthermore, in time series forecasting one should not randomly sample out the data as it is not a classification problem and it will lose the temporal pattern of the time series. Hope it solves your query..!

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.