Qlib Alpha158 Linear Model Backtest on CSI 500
Summary
This configuration defines a Qlib workflow that trains an ordinary least squares linear model on Alpha158 features for CSI 500 stocks. The data spans 2008 through mid-2020, with training through 2014, validation in 2015–2016, and testing from 2017 onward. Feature values are robustly normalized and missing values are filled; labels are cross-sectionally rank-normalized after rows with missing labels are dropped.
The workflow records signal analysis and a portfolio backtest. Its strategy holds the top 50 ranked names and allows five positions to be dropped, using the CSI 500 index as benchmark. The backtest uses closing prices and specifies transaction costs and a price-limit threshold. These are experiment settings rather than reported findings: the document provides no performance results, validation conclusions, or discussion of data quality and implementation assumptions. The chosen dates and trading assumptions therefore describe one evaluation setup, not evidence that the model is profitable or robust.
Key ideas
- The workflow applies an ordinary least squares model to Alpha158 features for CSI 500 stocks.
- The data is split into training, validation, and test periods, with testing beginning in 2017.
- Features receive robust normalization and missing-value filling, while labels are cross-sectionally rank-normalized.
- The portfolio strategy selects 50 names and permits five holdings to be dropped.
- The configuration specifies benchmark, closing-price execution, transaction costs, and price-limit assumptions but gives no backtest outcomes.
Tags
Full text
# workflow_config_linear_Alpha158_csi500.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors:
- class: RobustZScoreNorm
kwargs:
fields_group: feature
clip_outlier: true
- class: Fillna
kwargs:
fields_group: feature
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LinearModel
module_path: qlib.contrib.model.linear
kwargs:
estimator: ols
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: True
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.