Qlib Alpha158 Linear Model and Top-K Dropout Backtest Configuration
Summary
This Qlib workflow configuration specifies a linear ordinary least squares model using the Alpha158 feature handler for the CSI 300 universe. The data span begins in 2008 and ends in 2020; the training segment runs through 2014, validation covers 2015–2016, and testing covers 2017 through mid-2020. Feature processing applies robust z-score normalization and fills missing values, while labels are cross-sectionally rank normalized after rows with missing labels are dropped.
For portfolio analysis, the configuration uses a top-k dropout strategy that holds up to 50 names and replaces five at a time. Its backtest uses the Shanghai 300 benchmark, close prices, stated transaction-cost assumptions, a minimum fee, and a price-limit threshold. Signal and portfolio analysis records are enabled, including long-short signal analysis. This is an experimental setup, not a report of results: it contains no measured returns, risk statistics, or confirmation that the data, model, and execution assumptions avoid lookahead or other biases.
Key ideas
- The workflow trains an OLS linear model on Alpha158 features for the CSI 300 universe.
- The data are divided into training, validation, and test periods, with feature normalization and missing-value handling.
- Portfolio simulation uses a top-50 strategy that drops and replaces five holdings at a time.
- The backtest specifies a benchmark, close-price execution, costs, and a price-limit threshold.
- The configuration reports no performance results or validation of its assumptions.
Tags
Full text
# workflow_config_linear_Alpha158_multi_pass_bt.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors:
- class: RobustZScoreNorm
kwargs:
fields_group: feature
clip_outlier: true
- class: Fillna
kwargs:
fields_group: feature
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal:
- <MODEL>
- <DATASET>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LinearModel
module_path: qlib.contrib.model.linear
kwargs:
estimator: ols
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: True
ann_scaler: 252
- class: MultiPassPortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.