Skip to content
All library documents

Qlib CatBoost Alpha158 Workflow for CSI 300 Ranking

Code Qlib

Summary

This configuration describes a Qlib machine-learning workflow that trains a CatBoost regression model on Alpha158 features for CSI 300 instruments. It defines separate training, validation, and test periods, then records signal analysis and portfolio analysis. The portfolio simulation uses a top-k dropout strategy that holds a ranked set of stocks and replaces a limited number of positions as rankings change.

The backtest specifies a benchmark, account size, closing-price execution, transaction costs, minimum fees, and a price-limit threshold. Model settings include the RMSE objective, learning rate, subsampling, tree depth, and growth policy. These details make the file useful as an experiment template, but it contains no reported signal quality, returns, or risk results. Its conclusions would depend on data quality, feature construction, trading assumptions, and whether the time split prevents leakage; the configuration alone does not validate the strategy.

Key ideas

  • The workflow applies CatBoost regression to Alpha158 features for CSI 300 stocks.
  • It separates historical data into training, validation, and test intervals.
  • Portfolio analysis uses a top-k dropout ranking strategy with turnover controls.
  • The simulation specifies a benchmark, closing-price fills, fees, and a price-limit threshold.
  • The configuration reports no performance results and does not by itself establish strategy validity.

Tags

Full text
# workflow_config_catboost_Alpha158.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: CatBoostModel
        module_path: qlib.contrib.model.catboost_model
        kwargs:
            loss: RMSE
            learning_rate: 0.0421
            subsample: 0.8789
            max_depth: 6
            num_leaves: 100
            thread_count: 20
            grow_policy: Lossguide
            bootstrap_type: Poisson
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.