Qlib XGBoost Workflow for CSI 300 Stock Ranking
Summary
This configuration describes a Qlib workflow that trains an XGBoost model on China A-share data for the CSI 300 universe, using Alpha360 features. Its label is a forward close-to-close return over the next interval. Training and validation use earlier periods, while a later period is reserved for testing and portfolio analysis. The learning pipeline drops missing labels and applies cross-sectional rank normalization to labels.
Predicted signals feed a TopkDropoutStrategy that holds a ranked subset and replaces a limited number of holdings as rankings change. The backtest specifies a benchmark, closing-price deal assumptions, transaction costs, and a price-limit threshold; signal and portfolio analysis records are also configured. This is an experimental recipe, not evidence of predictive skill or profitability. The excerpt gives no performance results and does not establish that its splits, execution assumptions, or configured parameters remain suitable across markets or periods.
Key ideas
- The workflow uses Alpha360 features and an XGBoost model for CSI 300 instruments.
- The target is a forward return derived from closing prices.
- Training, validation, and testing are assigned to separate historical periods.
- A ranked top holdings strategy replaces a limited number of positions as signals change.
- The backtest includes benchmark, price-limit, and transaction-cost assumptions but reports no results.
Tags
Full text
# workflow_config_xgboost_Alpha360.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: []
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: XGBModel
module_path: qlib.contrib.model.xgboost
kwargs:
eval_metric: rmse
colsample_bytree: 0.8879
eta: 0.0421
max_depth: 8
n_estimators: 647
subsample: 0.8789
nthread: 20
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha360
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.