CatBoost Alpha360 Configuration for a CSI 500 Stock Portfolio Backtest
Summary
This configuration describes a Qlib workflow that trains a CatBoost model on Alpha360 features for China’s CSI 500 universe. Its label is a forward close-to-close return, normalized cross-sectionally after missing labels are dropped. The data is divided into training, validation, and test periods, with model fitting restricted to the training interval. CatBoost is configured for RMSE loss with specified learning and tree-growth settings.
For portfolio analysis, predicted scores feed a TopkDropoutStrategy that holds a ranked group of stocks and replaces some holdings as ranks change. The backtest uses closing prices, a benchmark, transaction costs, a minimum commission, and a price-limit threshold. Signal and portfolio analysis records are enabled. The file supplies experimental design and parameter choices, but no realized performance, robustness checks, or comparison with alternative models. Its results would depend on the underlying data, execution assumptions, and choices such as the label horizon and rebalance behavior; those settings alone do not establish predictive value.
Key ideas
- The workflow applies CatBoost to Alpha360 features for the CSI 500 stock universe.
- The target is a forward close-price return processed with cross-sectional rank normalization.
- Separate training, validation, and test intervals are configured.
- Portfolio construction uses a top-ranked holdings strategy with periodic dropout and replacement.
- The backtest specifies transaction costs, a benchmark, closing-price fills, and a price-limit assumption.
- The configuration reports no model performance or robustness evidence.
Tags
Full text
# workflow_config_catboost_Alpha360_csi500.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: []
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: CatBoostModel
module_path: qlib.contrib.model.catboost_model
kwargs:
loss: RMSE
learning_rate: 0.0421
subsample: 0.8789
max_depth: 6
num_leaves: 100
thread_count: 20
grow_policy: Lossguide
bootstrap_type: Poisson
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha360
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.