LightGBM Alpha360 Workflow for CSI 500 Stock Ranking
Summary
This configuration defines a Qlib workflow for ranking CSI 500 equities with a LightGBM model and Alpha360 features. It uses cross-sectional rank normalization for labels and trains on data from 2008 through 2014, validates on 2015–2016, and evaluates a later test period beginning in 2017. The target is based on a forward close-price return over the next interval specified in the label expression.
Portfolio construction uses a top-k dropout strategy, holding 50 names and replacing up to five, with the Shanghai Shenzhen 300? No: the stated benchmark is SH000905, the CSI 500 index. Backtest settings include close-price dealing, a 9.5% limit threshold, and stated open, close, and minimum transaction costs. The file records signal analysis and portfolio analysis but supplies no resulting metrics. Performance conclusions cannot be drawn from the configuration alone, and its historical windows, assumptions, and market-specific settings limit generalization.
Key ideas
- The workflow pairs Qlib Alpha360 features with a LightGBM regression model for CSI 500 equities.
- Labels receive cross-sectional rank normalization after missing labels are dropped.
- Training, validation, and test periods are separated across the stated historical dates.
- A top-k dropout portfolio holds 50 stocks and allows five positions to be replaced.
- The configuration specifies trading limits and transaction costs but does not provide backtest outcomes.
Tags
Full text
# workflow_config_lightgbm_Alpha360_csi500.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi500
benchmark: &benchmark SH000905
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: []
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs:
loss: mse
colsample_bytree: 0.8879
learning_rate: 0.0421
subsample: 0.8789
lambda_l1: 205.6999
lambda_l2: 580.9768
max_depth: 8
num_leaves: 210
num_threads: 20
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha360
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.