Qlib CatBoost Alpha360 Configuration for CSI 300 Portfolio Backtesting
Summary
This configuration defines a Qlib workflow that trains a CatBoost model on Alpha360 features for CSI 300 constituents. It uses Chinese market data from 2008 through mid-2020, with training through 2014, validation over 2015–2016, and testing from 2017 onward. The label is a forward close-price return, and the model uses an RMSE objective with specified tree-growth and sampling settings.
For portfolio analysis, the configuration applies a top-k dropout strategy that holds 50 names and replaces five, then backtests against the CSI 300 benchmark. It specifies a 100-million account, closing prices for deals, a 9.5% limit threshold, and transaction costs. Signal and portfolio analysis records are enabled. This is an experimental setup, not a report of findings: it contains no metrics, plots, or evidence of predictive or portfolio performance. Results would depend on Qlib data, implementation details, trading assumptions, and potential biases in the evaluation.
Key ideas
- The workflow pairs Qlib's Alpha360 handler with a CatBoost model for CSI 300 stocks.
- Training, validation, and test periods are separated, with the test segment beginning in 2017.
- The portfolio strategy holds 50 top-ranked names and drops five positions as signals change.
- The backtest specifies a CSI 300 benchmark, closing-price fills, transaction costs, and an account size.
- The configuration provides no reported results, so strategy performance cannot be inferred from it.
Tags
Full text
# workflow_config_catboost_Alpha360.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: []
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: CatBoostModel
module_path: qlib.contrib.model.catboost_model
kwargs:
loss: RMSE
learning_rate: 0.0421
subsample: 0.8789
max_depth: 6
num_leaves: 100
thread_count: 20
grow_policy: Lossguide
bootstrap_type: Poisson
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha360
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.