Qlib Double-Ensemble Backtest Configuration for CSI 500
Summary
This configuration defines a Qlib equity-prediction experiment for China’s CSI 500 universe. It trains a double-ensemble model built on gradient boosting, using Alpha360 features and a forward close-to-close return label. Training and validation segments precede a later test segment, with cross-sectional label normalization and missing-label filtering. The model combines sample reweighting and feature selection across multiple submodels, with specified sampling, weighting, and boosting settings.
Portfolio evaluation uses a top-k dropout strategy that holds a selected group of stocks and replaces only some positions as signals change. The backtest specifies a benchmark, close-price execution, assumed transaction costs, minimum fees, and a daily price-limit threshold. Signal analysis and portfolio analysis records are enabled. These are experimental settings, not reported results: the file gives no performance metrics, feature definitions beyond the Alpha360 handler, or evidence that the setup generalizes beyond this universe and period. The return label and portfolio settings also make validation sensitive to timing and execution assumptions.
Key ideas
- The setup applies a double-ensemble gradient-boosting model to Alpha360 data for the CSI 500 universe.
- Its label uses a future close-price return, while training labels are filtered for missing values and cross-sectionally ranked.
- The model enables both sample reweighting and feature selection across several submodels.
- Portfolio simulation uses a top-k dropout strategy with a China equity benchmark and explicit cost and price-limit assumptions.
- The configuration specifies an experiment but provides no results, so it cannot establish profitability or robustness.
Tags
Full text
# workflow_config_doubleensemble_Alpha360_csi500.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: []
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: DEnsembleModel
module_path: qlib.contrib.model.double_ensemble
kwargs:
base_model: "gbm"
loss: mse
num_models: 6
enable_sr: True
enable_fs: True
alpha1: 1
alpha2: 1
bins_sr: 10
bins_fs: 5
decay: 0.5
sample_ratios:
- 0.8
- 0.7
- 0.6
- 0.5
- 0.4
sub_weights:
- 1
- 0.2
- 0.2
- 0.2
- 0.2
- 0.2
epochs: 136
colsample_bytree: 0.8879
learning_rate: 0.0421
subsample: 0.8789
lambda_l1: 205.6999
lambda_l2: 580.9768
max_depth: 8
num_leaves: 210
num_threads: 20
verbosity: -1
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha360
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.