Skip to content
All library documents

Qlib LightGBM Alpha158 Model and CSI 300 Top-K Backtest Configuration

Code Qlib

Summary

This configuration specifies a Qlib workflow for training a LightGBM model on the Alpha158 feature set for China’s CSI 300 universe. It defines training, validation, and test periods, along with model settings such as mean squared error loss, tree depth, learning rate, sampling, and regularization. The workflow records signal analysis and portfolio analysis.

The portfolio simulation uses a top-K dropout strategy that holds 50 names and drops five positions, with the CSI 300 index as benchmark. The backtest covers 2017 through mid-2020 and specifies closing-price execution, transaction costs, a minimum fee, and a limit threshold. These details describe an experimental setup rather than evidence of profitability: the document includes no model metrics, trading results, or comparisons. Its conclusions would depend on the data, label construction, execution assumptions, and controls for look-ahead bias, none of which are discussed in this configuration alone.

Key ideas

  • The workflow trains a LightGBM model using Qlib’s Alpha158 features on the CSI 300 universe.
  • It separates historical data into training, validation, and test periods.
  • Portfolio analysis uses a top-K dropout strategy against the CSI 300 benchmark.
  • The configuration specifies transaction costs and closing-price execution assumptions.
  • The file defines an experiment but does not report results or validate its assumptions.

Tags

Full text
# workflow_config_lightgbm_Alpha158.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: LGBModel
        module_path: qlib.contrib.model.gbdt
        kwargs:
            loss: mse
            colsample_bytree: 0.8879
            learning_rate: 0.2
            subsample: 0.8789
            lambda_l1: 205.6999
            lambda_l2: 580.9768
            max_depth: 8
            num_leaves: 210
            num_threads: 20
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.