Skip to content
All library documents

Qlib CatBoost Alpha158 Model and CSI 500 Backtest Configuration

Code Qlib

Summary

This configuration specifies a Qlib workflow for training a CatBoost model on Alpha158 features for the CSI 500 universe. The data spans 2008 through mid-2020, with training through 2014, validation during 2015–2016, and testing from 2017 onward. The model uses RMSE loss and sets learning, sampling, depth, tree growth, and bootstrap options. Signal analysis and portfolio analysis are included as workflow records.

The portfolio simulation uses a top-k dropout strategy that holds 50 names and replaces five, with the CSI 500 index as benchmark. It sets an account value, closing-price execution, transaction costs, minimum fees, and a price-limit threshold. These settings make the document useful as a reproducible experiment outline, but the configuration reports no results. It also does not describe the feature definitions, data-cleaning choices, model tuning process, or safeguards against leakage, so performance cannot be inferred from the setup alone.

Key ideas

  • The workflow trains a CatBoost model using Alpha158 features on the CSI 500 universe.
  • The dataset separates training, validation, and test periods across 2008–2020.
  • Portfolio analysis applies a top-k dropout strategy with specified turnover and trading costs.
  • The configuration includes benchmark, execution price, price limit, and fee assumptions.
  • No model or backtest outcomes are reported.

Tags

Full text
# workflow_config_catboost_Alpha158_csi500.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: CatBoostModel
        module_path: qlib.contrib.model.catboost_model
        kwargs:
            loss: RMSE
            learning_rate: 0.0421
            subsample: 0.8789
            max_depth: 6
            num_leaves: 100
            thread_count: 20
            grow_policy: Lossguide
            bootstrap_type: Poisson
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.