Rolling Model Adaptation for Non-Stationary Market Forecasts
Summary
The document motivates adapting forecasting models to changing market conditions: financial data distributions can shift over time, so models trained on earlier periods may lose predictive strength. It compares two approaches, RR and DDG-DA, using linear and LightGBM forecasting models on the Alpha158 dataset. The reported table includes correlation-based metrics, annualized return, information ratio, and maximum drawdown.
The evaluation uses a 20-period label horizon and rolling intervals of 20 trading days, with test periods spanning January 2017 through August 2020. Results are drawn from a crowd-sourced data version. The document notes that an alternative Yahoo data version lacks VWAP, leaving related factors as zeros; this creates a rank-deficient matrix that prevents a lower-level DDG-DA optimization from being solved. The comparison is limited to the stated dataset, models, and period, and does not establish that the results generalize to other markets or future conditions.
Key ideas
- Changing market data distributions can cause forecasting performance to decay out of sample.
- The table compares RR and DDG-DA with linear and LightGBM models on Alpha158.
- The evaluation uses a 20-period label horizon and rolling intervals of 20 trading days.
- The test results cover January 2017 through August 2020 and rely on crowd-sourced data.
- Missing VWAP in the Yahoo data version leads to zero-filled factors and an unsolvable DDG-DA optimization step.
Tags
Full text
# Introduction # Introduction Due to the non-stationary nature of the environment of the financial market, the data distribution may change in different periods, which makes the performance of models build on training data decays in the future test data. So adapting the forecasting models/strategies to market dynamics is very important to the model/strategies' performance. The table below shows the performances of different solutions on different forecasting models. ## Alpha158 Dataset Here is the [crowd sourced version of qlib data](data_collector/crowd_source/README.md): https://github.com/chenditc/investment_data/releases ```bash wget https://github.com/chenditc/investment_data/releases/latest/download/qlib_bin.tar.gz mkdir -p ~/.qlib/qlib_data/cn_data tar -zxvf qlib_bin.tar.gz -C ~/.qlib/qlib_data/cn_data --strip-components=2 rm -f qlib_bin.tar.gz ``` | Model Name | Dataset | IC | ICIR | Rank IC | Rank ICIR | Annualized Return | Information Ratio | Max Drawdown | |------------------|---------|------|------|---------|-----------|-------------------|-------------------|--------------| | RR[Linear] |Alpha158 |0.0945|0.5989|0.1069 |0.6495 |0.0857 |1.3682 |-0.0986 | | DDG-DA[Linear] |Alpha158 |0.0983|0.6157|0.1108 |0.6646 |0.0764 |1.1904 |-0.0769 | | RR[LightGBM] |Alpha158 |0.0816|0.5887|0.0912 |0.6263 |0.0771 |1.3196 |-0.0909 | | DDG-DA[LightGBM] |Alpha158 |0.0878|0.6185|0.0975 |0.6524 |0.1261 |2.0096 |-0.0744 | - The label horizon of the `Alpha158` dataset is set to 20. - The rolling time intervals are set to 20 trading days. - The test rolling periods are from January 2017 to August 2020. - The results are based on the crowd-sourced version. The Yahoo version of qlib data does not contain `VWAP`, so all related factors are missing and filled with 0, which leads to a rank-deficient matrix (a matrix does not have full rank) and makes lower-level optimization of DDG-DA can not be solved.
Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.