Using Validation Sets and K-Fold Cross-Validation for Forecast Models
Summary
The article explains how cross-validation can estimate a model’s out-of-sample prediction error and help choose its flexibility, using a FTSE 100 forecasting example. Predictors are lagged daily prices or returns, and the response is the next day’s value. The focus is model assessment rather than converting predictions into a complete trading strategy.
It compares a validation-set split with k-fold cross-validation. A single split is simple but its error estimate can vary with the random partition; holding out a large share of observations also leaves less data for fitting. In k-fold validation, each fold takes a turn as the holdout set, and the errors are averaged. The article describes five or ten folds as common choices and explains that leave-one-out validation lowers bias but can have high variance and computational cost. Its example uses polynomial regression and reports procedures for comparing error curves, but the excerpt gives no evidence of live or trading performance. Randomly shuffling financial observations can also disregard chronology, so the evaluation design must reflect time-series use.
Key ideas
- Cross-validation estimates prediction error on observations excluded from model fitting.
- A single validation split can produce unstable estimates and reduce the training sample.
- K-fold validation averages errors across repeated holdout folds.
- Leave-one-out validation uses nearly all observations for each fit but can be costly and high variance.
- Financial forecasts require evaluation splits that respect temporal ordering.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.