LSTM Alpha158 Model and CSI 300 Portfolio Backtest Configuration
Summary
This configuration describes a Chinese equity forecasting experiment using Qlib’s Alpha158 features and an LSTM model. It selects 20 features, applies robust feature normalization and missing-value filling, and ranks labels cross-sectionally. The prediction target is the two-day forward close-price return. The dataset uses a 20-step sequence and divides the specified history into training, validation, and test periods.
The model settings specify a two-layer LSTM, mean-squared-error loss, early stopping, and GPU use. Portfolio evaluation applies a top-k dropout strategy that holds up to 50 stocks and replaces five positions at a time, with transaction costs and a price-limit threshold. Signal and portfolio analysis records are included. The document provides experiment settings rather than results: it reports no predictive accuracy, benchmark comparison, or realized return. Conclusions would depend on data quality, execution assumptions, and validation beyond the stated sample periods.
Key ideas
- The experiment forecasts a two-day forward close-price return for CSI 300 constituents.
- It uses 20 selected Alpha158 features with robust normalization and missing-value filling.
- The LSTM receives sequences of 20 time steps and is trained with mean-squared-error loss.
- A top-k dropout strategy evaluates portfolio signals while accounting for specified costs and price limits.
- The configuration contains no performance results, so it does not establish that the model is profitable.
Tags
Full text
# workflow_config_lstm_Alpha158.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors:
- class: FilterCol
kwargs:
fields_group: feature
col_list: ["RESI5", "WVMA5", "RSQR5", "KLEN", "RSQR10", "CORR5", "CORD5", "CORR10",
"ROC60", "RESI10", "VSTD5", "RSQR60", "CORR60", "WVMA60", "STD5",
"RSQR20", "CORD60", "CORD10", "CORR20", "KLOW"
]
- class: RobustZScoreNorm
kwargs:
fields_group: feature
clip_outlier: true
- class: Fillna
kwargs:
fields_group: feature
learn_processors:
- class: DropnaLabel
- class: CSRankNorm
kwargs:
fields_group: label
label: ["Ref($close, -2) / Ref($close, -1) - 1"]
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: LSTM
module_path: qlib.contrib.model.pytorch_lstm_ts
kwargs:
d_feat: 20
hidden_size: 64
num_layers: 2
dropout: 0.0
n_epochs: 200
lr: 1e-3
early_stop: 10
batch_size: 800
metric: loss
loss: mse
n_jobs: 20
GPU: 0
dataset:
class: TSDatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
step_len: 20
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.