Building and Backtesting a CSI 300 Alpha101 LightGBM Strategy
Summary
This notebook outlines a machine learning workflow for daily CSI 300 constituent stocks. It loads historical bars and changing index membership filters, constructs an Alpha101 dataset, and divides the sample into training, validation, and test periods. The learning pipeline drops missing labels and normalizes labels cross-sectionally with z-scores before preparing and caching the dataset.
A LightGBM model is trained on the prepared data, used to generate test-period signals, and evaluated through dataset signal performance analysis. The workflow then feeds those signals into an equity backtesting engine with a top-k portfolio rule, turnover through dropping holdings, and a holding threshold, and compares performance with the index benchmark. The notebook includes charts but states no numerical results in the text. The sample ends before the test period used in the backtest, and the example does not discuss transaction costs, slippage, survivorship bias, or robustness across alternative periods and settings.
Key ideas
- The workflow builds an Alpha101 dataset from daily CSI 300 constituent stock bars and membership filters.
- It separates data into training, validation, and test periods, then processes labels by dropping missing values and applying cross-sectional z-score normalization.
- A LightGBM model is fitted to the dataset and produces signals for the test segment.
- The signals drive a top-k equity strategy in a backtesting engine and are compared with the CSI 300 benchmark.
- The text provides no numerical performance results or discussion of transaction costs and slippage.
Tags
Full text
# 准备数据
# 准备数据
```python
# 过滤Alphalens的warning
import warnings
warnings.filterwarnings("ignore", category=FutureWarning)
```
```python
# 加载模块
import polars as pl
from vnpy.trader.constant import Interval
from vnpy.alpha import AlphaLab
```
```python
# 创建数据中心
lab: AlphaLab = AlphaLab("./lab/csi300")
```
```python
# 设置任务参数
name = "300_lgb"
index_symbol: str = "000300.SSE"
start: str = "2008-01-01"
end: str = "2023-12-31"
interval: Interval = Interval.DAILY
extended_days: int = 100
```
```python
# 加载所有成分股代码
component_symbols: list[str] = lab.load_component_symbols(index_symbol, start, end)
```
# 特征计算
```python
# 加载模块
from functools import partial
from vnpy.trader.constant import Interval
from vnpy.alpha.dataset import (
AlphaDataset,
process_drop_na,
process_cs_norm
)
from vnpy.alpha.dataset.datasets.alpha_101 import Alpha101
```
```python
# 加载成分股数据
df: pl.DataFrame = lab.load_bar_df(component_symbols, interval, start, end, extended_days)
```
```python
df
```
```python
# 创建数据集对象
dataset: AlphaDataset = Alpha101(
df,
train_period = ("2008-01-01", "2014-12-31"),
valid_period = ("2015-01-01", "2016-12-31"),
test_period = ("2017-01-01", "2020-8-31"),
)
```
```python
# 添加数据预处理器
dataset.add_processor("learn", partial(process_drop_na, names=["label"]))
dataset.add_processor("learn", partial(process_cs_norm, names=["label"], method="zscore"))
```
```python
# 收集指数成分过滤器
filters: dict[str, list[str]] = lab.load_component_filters(index_symbol, start, end)
```
```python
# 准备特征和标签数据
dataset.prepare_data(filters, max_workers=6)
```
```python
# 特征表现分析
dataset.show_feature_performance("alpha36")
```
```python
# 数据预处理
dataset.process_data()
```
```python
# 保存到文件缓存
lab.save_dataset(name, dataset)
```
# 模型训练
```python
# 加载模块
import numpy as np
from vnpy.alpha import Segment, AlphaDataset, AlphaModel
from vnpy.alpha.model.models.lgb_model import LgbModel
```
```python
# 从文件缓存加载
dataset: AlphaDataset = lab.load_dataset(name)
```
```python
# 创建模型对象
model: AlphaModel = LgbModel(seed=42)
```
```python
# 使用数据集训练模型
model.fit(dataset)
```
```python
# 查看模型细节
model.detail()
```
```python
# 保存模型
lab.save_model(name, model)
```
# 预测信号
```python
model: AlphaModel = lab.load_model(name)
```
```python
# 用模型在测试集上预测
pre: np.ndarray = model.predict(dataset, Segment.TEST)
# 加载测试集数据
df_t: pl.DataFrame = dataset.fetch_infer(Segment.TEST)
# 合并预测信号列
df_t = df_t.with_columns(pl.Series(pre).alias("signal"))
# 提取信号数据
signal: pl.DataFrame = df_t["datetime", "vt_symbol", "signal"]
```
```python
# 检查信号绩效
dataset.show_signal_performance(signal)
```
```python
# 保存信号数据
lab.save_signal(name, signal)
```
# 策略回测
```python
# 加载模块
import importlib
from datetime import datetime
from vnpy.alpha.strategy import BacktestingEngine
import vnpy.alpha.strategy.strategies.equity_demo_strategy as equity_demo_strategy
```
```python
# 重载策略类
importlib.reload(equity_demo_strategy)
EquityDemoStrategy = equity_demo_strategy.EquityDemoStrategy
```
```python
# 从文件加载信号数据
signal = lab.load_signal(name)
```
```python
# 创建回测引擎对象
engine = BacktestingEngine(lab)
# 设置回测参数
engine.set_parameters(
vt_symbols=component_symbols,
interval=Interval.DAILY,
start=datetime(2017, 1, 1),
end=datetime(2020, 8, 1),
capital=100000000
)
# 添加策略实例
setting = {"top_k": 30, "n_drop": 3, "hold_thresh": 3}
engine.add_strategy(EquityDemoStrategy, setting, signal)
```
```python
# 执行回测任务
engine.load_data()
engine.run_backtesting()
engine.calculate_result()
engine.calculate_statistics()
engine.show_chart()
```
```python
# 显示超额收益分析结果
engine.show_performance(benchmark_symbol=index_symbol)
```
```python
```

Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.