Skip to content
All library documents

Linear Alpha158 Modeling and CSI 300 Portfolio Backtesting

Code Qlib

Summary

This configuration defines an equity prediction workflow for China’s CSI 300 universe. It applies robust feature normalization and missing-value filling, drops labels with missing values, and ranks labels cross-sectionally. An ordinary least squares linear model is trained on an early period, validated on a later period, and evaluated on a held-out period.

For portfolio analysis, the workflow uses a top-ranked holdings strategy that replaces a limited number of positions as rankings change. It specifies a benchmark, closing-price execution, transaction costs, a minimum fee, and a price-limit threshold. Signal analysis and portfolio analysis are recorded, giving a framework for evaluating predictions and simulated performance. The document provides configuration rather than results: it reports no returns, risk statistics, or evidence that the strategy is profitable. Conclusions would also depend on data quality, the chosen dates, and how the simulation models trading constraints and costs.

Key ideas

  • The workflow uses Alpha158 features to predict cross-sectional stock outcomes in the CSI 300 universe.
  • An ordinary least squares model is trained, validated, and tested on separate time periods.
  • Feature values are normalized robustly and missing feature values are filled.
  • The portfolio strategy holds highly ranked stocks and replaces a subset as rankings change.
  • The backtest includes a benchmark, transaction costs, closing prices, and a price-limit constraint.

Tags

Full text
# workflow_config_linear_Alpha158.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
    infer_processors:
        - class: RobustZScoreNorm
          kwargs:
              fields_group: feature
              clip_outlier: true
        - class: Fillna
          kwargs:
              fields_group: feature
    learn_processors:
        - class: DropnaLabel
        - class: CSRankNorm
          kwargs:
              fields_group: label
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: LinearModel
        module_path: qlib.contrib.model.linear
        kwargs:
            estimator: ols
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: True
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.