LightGBM Alpha158 Model with Daily Labels and Minute Features
Summary
This configuration specifies a Chinese equities prediction and portfolio backtest using Qlib, LightGBM, and the Alpha158 feature handler. It pairs daily labels with one-minute features, resampling the minute data at 14:56. The listed data span begins in 2008 and ends in 2020, with fitting through 2014, validation over 2015–2016, and testing from 2017 through the stated end date.
The portfolio uses a top-k dropout strategy selecting 50 names and replacing up to five, with the CSI 300 as benchmark. The setup specifies closing-price deal assumptions, transaction costs, a limit threshold, and a fixed account value. LightGBM is configured for mean squared error with tree and regularization parameters. Signal, signal-analysis, and portfolio-analysis records are requested. This is an experiment specification, not a report of results: it gives no predictive metrics, returns, risk statistics, or evidence of out-of-sample success. Practical conclusions would depend on data quality and whether the configured costs and execution assumptions reflect actual trading.
Key ideas
- The model uses Alpha158 features with one-minute inputs resampled near the end of the trading day and daily labels.
- The configuration divides the stated history into training, validation, and test periods.
- LightGBM is trained with mean squared error and specified tree and regularization settings.
- The portfolio backtest applies a top-k dropout strategy against the CSI 300 benchmark.
- The configuration defines transaction cost and execution assumptions but reports no backtest results.
Tags
Full text
# workflow_config_lightgbm_Alpha158_multi_freq.yaml
```yaml
qlib_init:
provider_uri:
day: "~/.qlib/qlib_data/cn_data"
1min: "~/.qlib/qlib_data/cn_data_1min"
region: cn
dataset_cache: null
maxtasksperchild: 1
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
# 1min closing time is 15:00:00
end_time: "2020-08-01 15:00:00"
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
freq:
label: day
feature: 1min
# with label as reference
inst_processors:
feature:
- class: Resample1minProcessor
module_path: features_sample.py
trusted: true
kwargs:
hour: 14
minute: 56
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy.strategy
kwargs:
topk: 50
n_drop: 5
signal: <PRED>
backtest:
verbose: False
limit_threshold: 0.095
account: 100000000
benchmark: *benchmark
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LGBModel
module_path: qlib.contrib.model.gbdt
kwargs:
loss: mse
colsample_bytree: 0.8879
learning_rate: 0.2
subsample: 0.8789
lambda_l1: 205.6999
lambda_l2: 580.9768
max_depth: 8
num_leaves: 210
num_threads: 20
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs: {}
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.