Skip to content
All library documents

Qlib Enhanced Indexing with LightGBM on CSI 300 Data

Code Qlib

Summary

This configuration describes a Qlib workflow for training a LightGBM model and using its forecasts in an enhanced-indexing portfolio strategy. It specifies CSI 300 instruments and Shanghai’s CSI 300 index as the benchmark, with Alpha158 features and historical data spanning 2008 to 2020. The model is trained on the earlier period, validated on 2015–2016, and evaluated on 2017 through mid-2020.

Portfolio analysis uses a separate risk model and sets an initial account value, closing-price execution, transaction costs, and a price-limit threshold. Signal, signal-analysis, and portfolio-analysis records are enabled, so the setup connects prediction generation to a benchmark-relative backtest. The file provides configuration choices rather than reported findings: it gives no performance metrics, risk-model details, or explanation of how enhanced indexing constructs holdings. The placeholders for the model and dataset also need to be resolved in an actual run.

Key ideas

  • The workflow trains a LightGBM model using Alpha158 features on CSI 300 instruments.
  • The training, validation, and test segments divide the historical period chronologically.
  • An enhanced-indexing strategy uses a risk-model directory and the CSI 300 index benchmark.
  • The backtest configuration accounts for closing-price execution, transaction costs, and a price-limit threshold.
  • The configuration specifies analysis records but reports no backtest results.

Tags

Full text
# config_enhanced_indexing.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
port_analysis_config: &port_analysis_config
    strategy:
        class: EnhancedIndexingStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            model: <MODEL>
            dataset: <DATASET>
            riskmodel_root: ./riskdata
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: LGBModel
        module_path: qlib.contrib.model.gbdt
        kwargs:
            loss: mse
            colsample_bytree: 0.8879
            learning_rate: 0.2
            subsample: 0.8789
            lambda_l1: 205.6999
            lambda_l2: 580.9768
            max_depth: 8
            num_leaves: 210
            num_threads: 20
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record:
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs:
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs:
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs:
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.