Qlib CatBoost Alpha158 Workflow for CSI 300 Ranking
Summary
This configuration describes a Qlib machine-learning workflow that trains a CatBoost regression model on Alpha158 features for CSI 300 instruments. It defines separate training, validation, and test periods, then records signal analysis and portfolio analysis. The portfolio simulation uses a top-k dropout strategy that holds a ranked set of stocks and replaces a limited number of positions as rankings change.
The backtest specifies a benchmark, account size, closing-price execution, transaction costs, minimum fees, and a price-limit threshold. Model settings include the RMSE objective, learning rate, subsampling, tree depth, and growth policy. These details make the file useful as an experiment template, but it contains no reported signal quality, returns, or risk results. Its conclusions would depend on data quality, feature construction, trading assumptions, and whether the time split prevents leakage; the configuration alone does not validate the strategy.
Key ideas
- The workflow applies CatBoost regression to Alpha158 features for CSI 300 stocks.
- It separates historical data into training, validation, and test intervals.
- Portfolio analysis uses a top-k dropout ranking strategy with turnover controls.
- The simulation specifies a benchmark, closing-price fills, fees, and a price-limit threshold.
- The configuration reports no performance results and does not by itself establish strategy validity.
Tags
Full text
# workflow_config_catboost_Alpha158.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: CatBoostModel
module_path: qlib.contrib.model.catboost_model
kwargs:
loss: RMSE
learning_rate: 0.0421
subsample: 0.8789
max_depth: 6
num_leaves: 100
thread_count: 20
grow_policy: Lossguide
bootstrap_type: Poisson
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.