Partitioning Time Series for Forecasts with Future Labels
Summary
The document discusses how to evaluate a market forecasting model when each observation at time T is paired with an outcome measured n days later. Randomly assigning calendar years to training and testing sets can create boundary problems because future labels near a split may fall in the other set. Dropping boundary observations can also remove periods that matter, such as year-end events.
The response recommends evaluating model estimates against realized outcomes over rolling periods, with the measurement windows aligned; volatility estimates and realized volatility are given as an example. It cautions that fitting on an old period and testing much later may be uninformative when the model is dynamic. The answer does not provide a formal split algorithm or a detailed treatment of leakage and dependence between overlapping forecast horizons. It suggests handling recurring calendar effects, such as earnings announcements, through model adjustments or suitable evaluation windows.
Key ideas
- Future-dated labels can cross training and test boundaries when observations are split by calendar year.
- Rolling evaluation should align the model’s forecast horizon with the period used to measure the realized outcome.
- A model whose behavior changes over time may not be meaningfully assessed by training in one distant era and testing in another.
- Calendar events inside the forecast horizon may require explicit model adjustments.
- The response offers broad evaluation guidance rather than a complete procedure for preventing time-series leakage.
Tags
Full text
# Proper Data Partitioning For Building a Forecasting Model # Proper Data Partitioning For Building a Forecasting Model Goal: A team and I are looking to build a model that performs a predictive action for the state of the market on day `T + n`, using the data at hand on day `T`. To build this model, I'm using an EOD market data source going back until the early 1990s. Moreover, we are looking to separate the dataset into training and testing subsets, in order to optimize a model parameter, then to eventually test its performance in the out-of-sample subset. Problem: The initial attempt is to split the dataset by calendar year, and randomly assign each year into either the testing or training set. However, the following issues have been voiced: - Since our model is attempting to predict `n` days into the future (we can assume `n` is less than, say 15 or 20, but likely greater than 2-3), our training dataset needs to pull the market state from `n` days in the future in order to do our analysis. This would seem to indicate the we either need to pull `n` days from the testing dataset, or that we would need to drop the last `n` days from training set - Any given `n`-day span during the year may have some significance for our analysis, whether it's an earnings report, or otherwise. In particular, the `n`-day window that ends the calendar year is considered an important period for our analysis, so dropping these `n` border points is not ideal (and also might result in systematic model bias) Question: Is there a proper way to partition training and testing datasets, given that our analysis requires that we use a the datapoint that occurs `n` days in the future? ## Answer by Chris (score -1) https://quant.stackexchange.com/a/49942 For situations like this, I've typically considered the dataset as a whole and simply plotted, or in some other way evaluated, a model estimate relative to an actual value over a rolling period. For instance in the case volatility modeling, compare model-estimated (via ARCH/GARCH, implied vol, etc) vol to realized vol on a rolling basis, making sure the calculation periods coincide. Insofar as your model is dynamic, it doesn't make sense to fit the model on data from 1996 and then test it on data from 2007, unless you expect it to work. It's also going to be problematic to approach it as you describe given the +/-n day look ahead/back around beginning and end of period (year). Regarding your second bullet, I'm not clear why you'd drop `n` border points. The issue you mention about earnings announcements falling within your n-day window will likely be a limitation of your model though unless you make some kind of accommodation. For instance, to take it back to vol modeling again, you see spikes in vol at both the beginning and end of days and the beginning and end of weeks. It's easy enough to get around this by considering something akin to a seasonality adjustment or to simply look over long enough periods that the spikes even out. You'd probably want to make some other accommodation to deal with this in your case (eg, earnings announcements are public knowledge well beforehand, make some adjustment to your model to work around them).
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.