Configuring an XGBoost Alpha158 Portfolio Backtest
Summary
This configuration specifies a Chinese equity research workflow using Qlib, an XGBoost model, and the Alpha158 feature handler. The dataset is divided chronologically into training, validation, and test periods, with the model fitted on the training interval and assessed on later data. The model configuration includes a root mean squared error evaluation metric and settings for tree depth, number of estimators, learning rate, feature sampling, and row subsampling.
For portfolio simulation, the workflow uses a top-k dropout strategy that holds 50 selected names and replaces five at a time. The backtest compares against the CSI 300 benchmark and specifies a close-price deal assumption, transaction costs, a minimum fee, and a price-limit threshold. Signal analysis and portfolio analysis are recorded. These settings describe an experiment, not its findings: the configuration reports no returns, risk statistics, robustness checks, or evidence that the model generalizes beyond the stated universe and dates.
Key ideas
- The workflow applies XGBoost to Alpha158 features for CSI 300 equities.
- Chronological train, validation, and test segments separate model fitting from later assessment.
- The portfolio strategy selects 50 names and drops five positions at a time.
- The simulated exchange includes transaction costs, a minimum fee, and a price-limit threshold.
- The configuration defines an experiment but provides no performance results or robustness evidence.
Tags
Full text
# workflow_config_xgboost_Alpha158.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: XGBModel
module_path: qlib.contrib.model.xgboost
kwargs:
eval_metric: rmse
colsample_bytree: 0.8879
eta: 0.0421
max_depth: 8
n_estimators: 647
subsample: 0.8789
nthread: 20
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.