Qlib Neural Network Configuration for CSI 300 Signal Backtesting
Summary
This configuration specifies a Qlib experiment for generating equity signals on the CSI 300 universe, using the Shanghai Shenzhen 300 index as its benchmark. It sets a historical data range and separates training, validation, and test periods. The data handler uses Alpha158 features, drops a VWAP field, fills missing feature values for inference, removes missing training rows and labels, and normalizes labels cross-sectionally.
The model is a PyTorch general neural network trained with mean squared error and Adam, with settings for learning rate, batch size, weight decay, and input dimension. Portfolio analysis applies a top-k dropout strategy that holds leading signals and replaces a subset over time, then backtests with closing prices, transaction costs, a minimum fee, and a limit threshold. Signal, long-short analysis, and portfolio records are requested. This is an experiment specification rather than a report of results; it includes a comment that model parameters may be incorrect, and provides no performance evidence or discussion of data leakage and other validation risks.
Key ideas
- The configuration defines a CSI 300 equity prediction experiment with separate training, validation, and test periods.
- Alpha158 features are processed by dropping a field and handling missing values differently for inference and learning.
- A PyTorch neural network is specified with mean squared error loss and the Adam optimizer.
- Portfolio evaluation uses a top-k dropout strategy and includes closing-price execution costs and a limit threshold.
- The file records an experiment setup, not measured results, and flags that some model parameters may be wrong.
Tags
Full text
# workflow_config_mlp.yaml
```yaml
qlib_init:
provider_uri: "~/.qlib/qlib_data/cn_data"
region: cn
market: &market csi300
benchmark: &benchmark SH000300
data_handler_config: &data_handler_config
start_time: 2008-01-01
end_time: 2020-08-01
fit_start_time: 2008-01-01
fit_end_time: 2014-12-31
instruments: *market
infer_processors: [
{
"class" : "DropCol",
"kwargs":{"col_list": ["VWAP0"]}
},
{
"class" : "CSZFillna",
"kwargs":{"fields_group": "feature"}
}
]
learn_processors: [
{
"class" : "DropCol",
"kwargs":{"col_list": ["VWAP0"]}
},
{
"class" : "DropnaProcessor",
"kwargs":{"fields_group": "feature"}
},
"DropnaLabel",
{
"class": "CSZScoreNorm",
"kwargs": {"fields_group": "label"}
}
]
process_type: "independent"
port_analysis_config: &port_analysis_config
strategy:
class: TopkDropoutStrategy
module_path: qlib.contrib.strategy
kwargs:
signal: <PRED>
topk: 50
n_drop: 5
backtest:
start_time: 2017-01-01
end_time: 2020-08-01
account: 100000000
benchmark: *benchmark
exchange_kwargs:
limit_threshold: 0.095
deal_price: close
open_cost: 0.0005
close_cost: 0.0015
min_cost: 5
task:
model:
class: GeneralPTNN
module_path: qlib.contrib.model.pytorch_general_nn
kwargs:
# FIXME: wrong parameters.
lr: 2e-3
batch_size: 8192
loss: mse
weight_decay: 0.0002
optimizer: adam
pt_model_uri: "qlib.contrib.model.pytorch_nn.Net"
pt_model_kwargs:
input_dim: 157
dataset:
class: DatasetH
module_path: qlib.data.dataset
kwargs:
handler:
class: Alpha158
module_path: qlib.contrib.data.handler
kwargs: *data_handler_config
segments:
train: [2008-01-01, 2014-12-31]
valid: [2015-01-01, 2016-12-31]
test: [2017-01-01, 2020-08-01]
record:
- class: SignalRecord
module_path: qlib.workflow.record_temp
kwargs:
model: <MODEL>
dataset: <DATASET>
- class: SigAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
ana_long_short: False
ann_scaler: 252
- class: PortAnaRecord
module_path: qlib.workflow.record_temp
kwargs:
config: *port_analysis_config
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.