Building and Backtesting a CSI 300 Alpha158 MLP Signal
Summary
This workflow demonstrates an equity prediction pipeline using CSI 300 constituent data, Alpha158 features, and a multilayer perceptron. It defines training, validation, and test periods, prepares constituent-filtered data, normalizes features using robust statistics fitted on the training period, fills missing feature values, and drops missing labels. It then fits the model, generates test-segment predictions, and saves the resulting signal.
The final stage passes the signal to a portfolio backtesting engine with a top-k selection and holding parameters, then requests performance and benchmark-relative analysis. The example shows an end-to-end process from data preparation through backtest, but reports no performance results. Its periods, model settings, and strategy parameters are illustrative choices; the document does not explain label construction, transaction costs, leakage checks, or whether the reported backtest is robust out of sample.
Key ideas
- The example uses CSI 300 constituents and Alpha158 features to build an equity prediction dataset.
- Feature normalization is fitted on the training interval, while missing feature values are filled and missing labels are dropped.
- An MLP is trained before producing predictions for the designated test segment.
- The predictions are converted into a signal and evaluated through a top-k portfolio backtest.
- No performance results or transaction cost assumptions are provided.
Tags
Full text
# 准备数据
# 准备数据
```python
# 过滤Alphalens的warning
import warnings
warnings.filterwarnings("ignore", category=FutureWarning)
```
```python
# 加载模块
import polars as pl
from vnpy.trader.constant import Interval
from vnpy.alpha import AlphaLab
```
```python
# 创建数据中心
lab: AlphaLab = AlphaLab("./lab/csi300")
```
```python
# 设置任务参数
name = "300_mlp"
index_symbol: str = "000300.SSE"
start: str = "2008-01-01"
end: str = "2023-12-31"
interval: Interval = Interval.DAILY
extended_days: int = 100
```
```python
# 加载所有成分股代码
component_symbols: list[str] = lab.load_component_symbols(index_symbol, start, end)
```
# 特征计算
```python
# 加载模块
from datetime import datetime
from functools import partial
from vnpy.trader.constant import Interval
from vnpy.alpha.dataset import (
AlphaDataset,
process_drop_na,
process_robust_zscore_norm,
process_fill_na,
process_cs_rank_norm,
to_datetime
)
from vnpy.alpha.dataset.datasets.alpha_158 import Alpha158
```
```python
# 加载成分股数据
df: pl.DataFrame = lab.load_bar_df(component_symbols, interval, start, end, extended_days)
```
```python
# 设置数据时间段
train_period: tuple[str, str] = ("2008-01-01", "2014-12-31")
valid_period: tuple[str, str] = ("2015-01-01", "2016-12-31")
test_period: tuple[str, str] = ("2017-01-01", "2020-8-31")
```
```python
# 创建数据集对象
dataset: AlphaDataset = Alpha158(
df,
train_period=train_period,
valid_period=valid_period,
test_period=test_period,
)
```
```python
# 添加数据预处理器
fit_start_time: datetime = to_datetime(train_period[0])
fit_end_time: datetime = to_datetime(train_period[1])
dataset.add_processor("infer", partial(process_robust_zscore_norm, fit_start_time=fit_start_time, fit_end_time=fit_end_time))
dataset.add_processor("infer", partial(process_fill_na, fill_value=0, fill_label=False))
dataset.add_processor("learn", partial(process_drop_na, names=["label"]))
dataset.add_processor("learn", partial(process_cs_rank_norm, names=["label"]))
```
```python
# 收集指数成分过滤器
filters: dict[str, list[str]] = lab.load_component_filters(index_symbol, start, end)
```
```python
# 准备特征和标签数据
dataset.prepare_data(filters, max_workers=3)
```
```python
# 数据预处理
dataset.process_data()
```
```python
# 特征表现分析
dataset.show_feature_performance("rsv_5")
```
```python
# 保存到文件缓存
lab.save_dataset(name, dataset)
```
# 模型训练
```python
# 加载模块
import numpy as np
from vnpy.alpha import Segment, AlphaDataset, AlphaModel
from vnpy.alpha.model.models.mlp_model import MlpModel
```
```python
# 从文件缓存加载
dataset: AlphaDataset = lab.load_dataset(name)
```
```python
# 创建模型对象
kwargs = {
"input_size": 158,
"hidden_sizes": (256,),
"lr": 0.002,
"optimizer": "adam",
"n_epochs": 8000,
"batch_size": 8192,
"weight_decay": 0.0002,
"seed": 42
}
model: AlphaModel = MlpModel(**kwargs)
```
```python
# 使用数据集训练模型
model.fit(dataset)
```
```python
# 查看模型细节
model.detail()
```
```python
# 保存模型
lab.save_model(name, model)
```
# 预测信号
```python
model: AlphaModel = lab.load_model(name)
```
```python
# 用模型在测试集上预测
pre: np.ndarray = model.predict(dataset, Segment.TEST)
# 加载测试集数据
df_t: pl.DataFrame = dataset.fetch_infer(Segment.TEST)
# 合并预测信号列
df_t = df_t.with_columns(pl.Series(pre).alias("signal"))
# 提取信号数据
signal: pl.DataFrame = df_t["datetime", "vt_symbol", "signal"]
```
```python
# 检查信号绩效
dataset.show_signal_performance(signal)
```
```python
# 保存信号数据
lab.save_signal(name, signal)
```
# 策略回测
```python
# 加载模块
import importlib
from datetime import datetime
from vnpy.alpha.strategy import BacktestingEngine
import vnpy.alpha.strategy.strategies.equity_demo_strategy as equity_demo_strategy
```
```python
# 重载策略类
importlib.reload(equity_demo_strategy)
EquityDemoStrategy = equity_demo_strategy.EquityDemoStrategy
```
```python
# 从文件加载信号数据
signal = lab.load_signal(name)
```
```python
# 创建回测引擎对象
engine = BacktestingEngine(lab)
# 设置回测参数
engine.set_parameters(
vt_symbols=component_symbols,
interval=Interval.DAILY,
start=datetime(2017, 1, 1),
end=datetime(2020, 8, 1),
capital=100000000
)
# 添加策略实例
setting = {"top_k": 30, "n_drop": 3, "hold_thresh": 3}
engine.add_strategy(EquityDemoStrategy, setting, signal)
```
```python
# 执行回测任务
engine.load_data()
engine.run_backtesting()
engine.calculate_result()
engine.calculate_statistics()
engine.show_chart()
```
```python
# 显示超额收益分析结果
engine.show_performance(benchmark_symbol=index_symbol)
```
```python
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.