Linear Alpha158 Modeling and CSI 300 Portfolio Backtesting
Summary
This configuration defines an equity prediction workflow for China’s CSI 300 universe. It applies robust feature normalization and missing-value filling, drops labels with missing values, and ranks labels cross-sectionally. An ordinary least squares linear model is trained on an early period, validated on a later period, and evaluated on a held-out period.
For portfolio analysis, the workflow uses a top-ranked holdings strategy that replaces a limited number of positions as rankings change. It specifies a benchmark, closing-price execution, transaction costs, a minimum fee, and a price-limit threshold. Signal analysis and portfolio analysis are recorded, giving a framework for evaluating predictions and simulated performance. The document provides configuration rather than results: it reports no returns, risk statistics, or evidence that the strategy is profitable. Conclusions would also depend on data quality, the chosen dates, and how the simulation models trading constraints and costs.
Key ideas
- The workflow uses Alpha158 features to predict cross-sectional stock outcomes in the CSI 300 universe.
- An ordinary least squares model is trained, validated, and tested on separate time periods.
- Feature values are normalized robustly and missing feature values are filled.
- The portfolio strategy holds highly ranked stocks and replaces a subset as rankings change.
- The backtest includes a benchmark, transaction costs, closing prices, and a price-limit constraint.
Tags
Full text
# workflow_config_linear_Alpha158.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors:
- class: RobustZScoreNorm
kwargs:
fields_group: feature
clip_outlier: true
- class: Fillna
kwargs:
fields_group: feature
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LinearModel
module_path: qlib.contrib.model.linear
kwargs:
estimator: ols
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: True
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.