Qlib Double-Ensemble Configuration for CSI 300 Stock Ranking
Summary
This configuration sets up a Qlib experiment that uses a double-ensemble model built from gradient-boosted trees to rank CSI 300 stocks. The dataset uses Alpha158 features and divides the history into training, validation, and test segments. Model settings enable sample reweighting and feature selection, combine multiple component models, and specify early stopping. The excerpt therefore provides a concrete experimental setup for researching stock-selection signals rather than an explanation of the model’s mechanics.
For portfolio analysis, the configuration applies a top-k dropout strategy: it holds a ranked group of stocks and replaces a smaller number as rankings change. It defines a benchmark, account size, transaction costs, a price limit, and closing-price execution assumptions for the backtest. Signal, signal-analysis, and portfolio-analysis records are requested. However, the file reports no resulting returns, risk statistics, or comparison with alternatives. Its conclusions cannot be assessed without running the experiment, and the specified historical sample and trading assumptions limit what a backtest could establish about future performance.
Key ideas
- The experiment uses a double-ensemble gradient-boosted model with Alpha158 features to rank CSI 300 stocks.
- Training, validation, and test periods are specified separately.
- The model configuration enables feature selection, sample reweighting, and early stopping.
- Portfolio analysis uses a top-k dropout strategy with a benchmark and explicit trading-cost assumptions.
- The configuration contains no reported performance results, so it does not establish profitability.
Tags
Full text
# workflow_config_doubleensemble_early_stop_Alpha158.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: DEnsembleModel
module_path: qlib.contrib.model.double_ensemble
kwargs:
base_model: "gbm"
loss: mse
num_models: 3
enable_sr: True
enable_fs: True
alpha1: 1
alpha2: 1
bins_sr: 10
bins_fs: 5
decay: 0.5
sample_ratios:
- 0.8
- 0.7
- 0.6
- 0.5
- 0.4
sub_weights:
- 1
- 1
- 1
epochs: 1000
early_stopping_rounds: 50
colsample_bytree: 0.8879
learning_rate: 0.2
subsample: 0.8789
lambda_l1: 205.6999
lambda_l2: 580.9768
max_depth: 8
num_leaves: 210
num_threads: 20
verbosity: -1
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.