Skip to content
All library documents

Building and Backtesting a CSI 300 Alpha158 MLP Signal

Notebook vn.py

Summary

This workflow demonstrates an equity prediction pipeline using CSI 300 constituent data, Alpha158 features, and a multilayer perceptron. It defines training, validation, and test periods, prepares constituent-filtered data, normalizes features using robust statistics fitted on the training period, fills missing feature values, and drops missing labels. It then fits the model, generates test-segment predictions, and saves the resulting signal.

The final stage passes the signal to a portfolio backtesting engine with a top-k selection and holding parameters, then requests performance and benchmark-relative analysis. The example shows an end-to-end process from data preparation through backtest, but reports no performance results. Its periods, model settings, and strategy parameters are illustrative choices; the document does not explain label construction, transaction costs, leakage checks, or whether the reported backtest is robust out of sample.

Key ideas

  • The example uses CSI 300 constituents and Alpha158 features to build an equity prediction dataset.
  • Feature normalization is fitted on the training interval, while missing feature values are filled and missing labels are dropped.
  • An MLP is trained before producing predictions for the designated test segment.
  • The predictions are converted into a signal and evaluated through a top-k portfolio backtest.
  • No performance results or transaction cost assumptions are provided.

Tags

Full text
# 准备数据


# 准备数据

```python
# 过滤Alphalens的warning
import warnings
warnings.filterwarnings("ignore", category=FutureWarning)
```

```python
# 加载模块
import polars as pl

from vnpy.trader.constant import Interval

from vnpy.alpha import AlphaLab
```

```python
# 创建数据中心
lab: AlphaLab = AlphaLab("./lab/csi300")
```

```python
# 设置任务参数
name = "300_mlp"
index_symbol: str = "000300.SSE"
start: str = "2008-01-01"
end: str = "2023-12-31"
interval: Interval = Interval.DAILY
extended_days: int = 100
```

```python
# 加载所有成分股代码
component_symbols: list[str] = lab.load_component_symbols(index_symbol, start, end)
```

# 特征计算

```python
# 加载模块
from datetime import datetime
from functools import partial

from vnpy.trader.constant import Interval

from vnpy.alpha.dataset import (
    AlphaDataset,
    process_drop_na,
    process_robust_zscore_norm,
    process_fill_na,
    process_cs_rank_norm,
    to_datetime
)
from vnpy.alpha.dataset.datasets.alpha_158 import Alpha158
```

```python
# 加载成分股数据
df: pl.DataFrame = lab.load_bar_df(component_symbols, interval, start, end, extended_days)
```

```python
# 设置数据时间段
train_period: tuple[str, str] = ("2008-01-01", "2014-12-31")
valid_period: tuple[str, str] = ("2015-01-01", "2016-12-31")
test_period: tuple[str, str] = ("2017-01-01", "2020-8-31")
```

```python
# 创建数据集对象
dataset: AlphaDataset = Alpha158(
    df,
    train_period=train_period,
    valid_period=valid_period,
    test_period=test_period,
)
```

```python
# 添加数据预处理器
fit_start_time: datetime = to_datetime(train_period[0])
fit_end_time: datetime = to_datetime(train_period[1])

dataset.add_processor("infer", partial(process_robust_zscore_norm, fit_start_time=fit_start_time, fit_end_time=fit_end_time))
dataset.add_processor("infer", partial(process_fill_na, fill_value=0, fill_label=False))

dataset.add_processor("learn", partial(process_drop_na, names=["label"]))
dataset.add_processor("learn", partial(process_cs_rank_norm, names=["label"]))
```

```python
# 收集指数成分过滤器
filters: dict[str, list[str]] = lab.load_component_filters(index_symbol, start, end)
```

```python
# 准备特征和标签数据
dataset.prepare_data(filters, max_workers=3)
```

```python
# 数据预处理
dataset.process_data()
```

```python
# 特征表现分析
dataset.show_feature_performance("rsv_5")
```

```python
# 保存到文件缓存
lab.save_dataset(name, dataset)
```

# 模型训练

```python
# 加载模块
import numpy as np

from vnpy.alpha import Segment, AlphaDataset, AlphaModel

from vnpy.alpha.model.models.mlp_model import MlpModel
```

```python
# 从文件缓存加载
dataset: AlphaDataset = lab.load_dataset(name)
```

```python
# 创建模型对象
kwargs = {
    "input_size": 158,
    "hidden_sizes": (256,),
    "lr": 0.002,
    "optimizer": "adam",
    "n_epochs": 8000,
    "batch_size": 8192,
    "weight_decay": 0.0002,
    "seed": 42
}

model: AlphaModel = MlpModel(**kwargs)
```

```python
# 使用数据集训练模型
model.fit(dataset)
```

```python
# 查看模型细节
model.detail()
```

```python
# 保存模型
lab.save_model(name, model)
```

# 预测信号

```python
model: AlphaModel = lab.load_model(name)
```

```python
# 用模型在测试集上预测
pre: np.ndarray = model.predict(dataset, Segment.TEST)

# 加载测试集数据
df_t: pl.DataFrame = dataset.fetch_infer(Segment.TEST)

# 合并预测信号列
df_t = df_t.with_columns(pl.Series(pre).alias("signal"))

# 提取信号数据
signal: pl.DataFrame = df_t["datetime", "vt_symbol", "signal"]
```

```python
# 检查信号绩效
dataset.show_signal_performance(signal)
```

```python
# 保存信号数据
lab.save_signal(name, signal)
```

# 策略回测

```python
# 加载模块
import importlib
from datetime import datetime

from vnpy.alpha.strategy import BacktestingEngine

import vnpy.alpha.strategy.strategies.equity_demo_strategy as equity_demo_strategy
```

```python
# 重载策略类
importlib.reload(equity_demo_strategy)
EquityDemoStrategy = equity_demo_strategy.EquityDemoStrategy
```

```python
# 从文件加载信号数据
signal = lab.load_signal(name)
```

```python
# 创建回测引擎对象
engine = BacktestingEngine(lab)

# 设置回测参数
engine.set_parameters(
    vt_symbols=component_symbols,
    interval=Interval.DAILY,
    start=datetime(2017, 1, 1),
    end=datetime(2020, 8, 1),
    capital=100000000
)

# 添加策略实例
setting = {"top_k": 30, "n_drop": 3, "hold_thresh": 3}
engine.add_strategy(EquityDemoStrategy, setting, signal)
```

```python
# 执行回测任务
engine.load_data()
engine.run_backtesting()
engine.calculate_result()
engine.calculate_statistics()
engine.show_chart()
```

```python
# 显示超额收益分析结果
engine.show_performance(benchmark_symbol=index_symbol)
```

```python

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.