Skip to content
All library documents

Preparing Historical CSI 300 Constituents and Daily Price Data

Notebook vn.py

Summary

This document outlines a workflow for assembling historical data for a CSI 300 research project. It downloads historical constituent information, retrieves the index membership for each trading date, converts vendor symbols into vn.py format, and saves the dated membership records. It then loads the constituent symbols for a chosen interval and downloads daily bars for those stocks and the index.

The workflow also adds contract settings for each constituent, including long and short commission rates, contract size, and price tick. These settings support later research or backtesting. The document provides code steps but no data-quality checks, backtest results, or discussion of how the data provider handles missing or revised historical records. Its usefulness therefore lies in the preparation process; users would need to verify symbol conversions, coverage, and time-zone handling for their own setup.

Key ideas

  • Historical index membership can be collected separately for each trading date.
  • The workflow saves dated CSI 300 constituent lists before downloading price bars.
  • Daily bars are requested for both constituent stocks and the index.
  • Contract settings include commission rates, size, and minimum price movement.
  • The document does not report validation checks or investment performance.

Tags

Full text
# 准备数据


# 准备数据

```python
# 加载模块
from datetime import datetime

from tqdm import tqdm
from xtquant import xtdata

from vnpy.trader.database import DB_TZ
from vnpy.trader.datafeed import get_datafeed
from vnpy.trader.constant import Exchange, Interval
from vnpy.trader.object import HistoryRequest

from vnpy.alpha import AlphaLab, logger
```

```python
# 设置下载参数
task_name = "csi300"
index_symbol = "000300.SSE"
xt_index_symbol = "000300.SH"

start_date = "20070101"
end_date = "20231231"

intervals = [
    Interval.DAILY,
]
```

```python
# 创建投研实验室
lab = AlphaLab(f"./lab/{task_name}")    # 指定数据文件夹
```

```python
# 初始化数据服务(这里配置使用的迅投研)
datafeed = get_datafeed()
datafeed.init()
```

```python
# 下载历史成分股信息
xtdata.download_sector_data()

xtdata.download_history_data("", "stocklistchange", "", "")
```

```python
# 查询交易日历
days = xtdata.get_trading_calendar(market="SZ", start_time=start_date, end_time=end_date)

# 轮询获取指数成本股
index_components = {}
end_datetime = datetime.strptime(end_date, "%Y%m%d")
for i in days:
    dt = datetime.strptime(i, "%Y%m%d")
    if dt > end_datetime:
        continue

    xt_symbols = xtdata.get_stock_list_in_sector(xt_index_symbol, i)

    vt_symbols: list = []
    for xt_symbol in xt_symbols:
        vt_symbol = xt_symbol.replace("SH", "SSE").replace("SZ", "SZSE")
        vt_symbols.append(vt_symbol)

    index_components[dt.strftime("%Y-%m-%d")] = vt_symbols

# 保存到数据中心
lab.save_component_data(index_symbol, index_components)
```

```python
# 加载指数成分股代码
component_symbols = lab.load_component_symbols(index_symbol, start_date, end_date)
```

```python
# 转换时间格式
start = datetime.strptime(start_date, "%Y%m%d")
start.replace(tzinfo=DB_TZ)

end = datetime.strptime(end_date, "%Y%m%d")
end.replace(tzinfo=DB_TZ)

# 除了成分股,还要下载指数数据
task_symbols = component_symbols + [index_symbol]

# 遍历下载数据
for vt_symbol in tqdm(task_symbols):
    symbol, exchange_str = vt_symbol.split(".")

    for interval in intervals:
        req = HistoryRequest(symbol, Exchange(exchange_str), start, end, interval)
        bars = datafeed.query_bar_history(req)

        if bars:
            lab.save_bar_data(bars)
        else:
            logger(interval, vt_symbol)
```

```python
# 添加回测参数配置
for vt_symbol in component_symbols:
    lab.add_contract_setting(
        vt_symbol,
        long_rate=5/10000,
        short_rate=10/10000,
        size=1,
        pricetick=0.0001,
    )
```

```python

```

Shown in full with attribution under the source's licence. Licence: MIT

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.