Skip to content
All library documents

Qlib LightGBM Alpha360 Model and CSI 300 Backtest Setup

Code Qlib

Summary

This configuration describes a Qlib workflow for training a LightGBM model on Alpha360 features and evaluating its predictions on CSI 300 constituents. The label compares closing prices across the next two reference periods. The data handler drops missing labels and rank-normalizes labels cross-sectionally; the dataset is divided into training, validation, and test periods, with fitting limited to the training span.

For portfolio evaluation, the configuration uses a top-k dropout strategy that holds up to 50 names and replaces up to five, with the CSI 300 index as benchmark. The backtest specifies closing-price execution, transaction costs, a minimum fee, and a price-limit threshold. Signal, signal-analysis, and portfolio-analysis records are enabled. These settings make the document useful as an example of an end-to-end equity modeling and backtesting workflow. It supplies configuration choices but no reported performance, robustness checks, or evidence that the model generalizes beyond the selected China-market data and dates.

Key ideas

  • The workflow trains a LightGBM model using Alpha360 features for CSI 300 equities.
  • The target is based on the ratio of closing prices across two future reference periods.
  • Training, validation, and test data occupy distinct date ranges, while model fitting uses the training period.
  • Portfolio evaluation uses a top-k dropout strategy and benchmarks results against the CSI 300 index.
  • The backtest specifies closing-price execution, transaction costs, and a price-limit threshold.

Tags

Full text
# workflow_config_lightgbm_Alpha360.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
    infer_processors: []
    learn_processors:
        - class: DropnaLabel
        - class: CSRankNorm
          kwargs:
              fields_group: label
    label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: LGBModel
        module_path: qlib.contrib.model.gbdt
        kwargs:
            loss: mse
            colsample_bytree: 0.8879
            learning_rate: 0.0421
            subsample: 0.8789
            lambda_l1: 205.6999
            lambda_l2: 580.9768
            max_depth: 8
            num_leaves: 210
            num_threads: 20
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha360
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.