Skip to content
All library documents

LightGBM Alpha158 Backtest Configuration for CSI 500 Stocks

Code Qlib

Summary

This configuration defines a Qlib workflow that trains a LightGBM model on Alpha158 features for the CSI 500 universe. It assigns data from 2008 through mid-2020 to training, validation, and test segments, with the fit period ending in 2014. The model uses mean squared error and specifies tree sampling, regularization, depth, leaf, and thread settings.

For portfolio analysis, it uses a top-k dropout strategy that holds 50 stocks and replaces up to 5, then backtests from 2017 through August 2020 against the CSI 500 benchmark. The setup specifies a one-hundred-million account, closing prices for trades, transaction costs, a minimum fee, and a price-limit threshold. Signal, long-short analysis, and portfolio analysis records are enabled. The file is a reproducible experiment specification, not a report of performance: it provides no results, comparisons, or evidence that the model or strategy is profitable. Outcomes also depend on the underlying data, execution assumptions, and implementation details.

Key ideas

  • The workflow applies LightGBM to Alpha158 features for CSI 500 stocks.
  • Training, validation, and test periods are explicitly separated, with model fitting ending before the test period.
  • The portfolio strategy holds 50 names and allows 5 positions to be dropped and replaced.
  • The backtest specifies benchmark, trading price, fees, and a price-limit assumption.
  • The configuration describes an experiment but reports no predictive or investment results.

Tags

Full text
# workflow_config_lightgbm_Alpha158_csi500.yaml


```yaml
qlib_init:
    provider_uri: "~/.qlib/qlib_data/cn_data"
    region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
    start_time: 2008-01-01
    end_time: 2020-08-01
    fit_start_time: 2008-01-01
    fit_end_time: 2014-12-31
    instruments: *market
port_analysis_config: &port_analysis_config
    strategy:
        class: TopkDropoutStrategy
        module_path: qlib.contrib.strategy
        kwargs:
            signal: <PRED>
            topk: 50
            n_drop: 5
    backtest:
        start_time: 2017-01-01
        end_time: 2020-08-01
        account: 100000000
        benchmark: *benchmark
        exchange_kwargs:
            limit_threshold: 0.095
            deal_price: close
            open_cost: 0.0005
            close_cost: 0.0015
            min_cost: 5
task:
    model:
        class: LGBModel
        module_path: qlib.contrib.model.gbdt
        kwargs:
            loss: mse
            colsample_bytree: 0.9
            learning_rate: 0.1
            subsample: 0.9
            lambda_l1: 205.6999
            lambda_l2: 580.9768
            max_depth: 8
            num_leaves: 250
            num_threads: 20
    dataset:
        class: DatasetH
        module_path: qlib.data.dataset
        kwargs:
            handler:
                class: Alpha158
                module_path: qlib.contrib.data.handler
                kwargs: *data_handler_config
            segments:
                train: [2008-01-01, 2014-12-31]
                valid: [2015-01-01, 2016-12-31]
                test: [2017-01-01, 2020-08-01]
    record: 
        - class: SignalRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            model: <MODEL>
            dataset: <DATASET>
        - class: SigAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            ana_long_short: False
            ann_scaler: 252
        - class: PortAnaRecord
          module_path: qlib.workflow.record_temp
          kwargs: 
            config: *port_analysis_config

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.