用于期限结构与持有收益研究的CME期货数据
文章 《交易机器学习》
总结
本文介绍CME期货数据集,包含多个产品组的小时线和日线、连续近月合约以及两个远期期限。文中说明数据覆盖范围、字段、载入选项,以及相关的每周CFTC持仓数据。期限结构支持期限结构分析和持有收益研究;重建和研究换月则需要单个合约记录。
材料主要是数据集指南,而非交易研究:它没有提供策略结果,也没有证据表明某种方法能够盈利。市场历史数据需要付费数据源和API密钥,完整刷新可能产生显著费用。连续合约已经过换月,因此仅凭它们无法判断换月是如何构造的。CFTC数据增加了交易者类别持仓信息,但数据按周发布且存在延迟。根据数据源许可,原始数据不得再分发。
核心观点
- 该数据集提供近月及两个远期期限的CME期货小时线和日线数据。
- 比较不同期限有助于进行期限结构和持有收益分析。
- 连续序列已经过换月,因此需要单个合约数据来重建换月过程。
- 每周CFTC报告提供交易者类别持仓快照,但发布存在延迟。
- 市场数据需要付费的API数据源,且数据源限制原始时间序列的再分发。
标签
全文
# CME Futures (Databento)
# CME Futures (Databento)
30 CME futures products — continuous front-month contracts plus the next
two tenors — for term-structure analysis, carry strategies, and the
`cme_futures` case study. Daily and hourly bars available.
## Dataset
- **Source**: Databento via the `GLBX.MDP3` CME dataset.
- **Coverage**: 2011-01-01 → present, hourly OHLCV with daily aggregation.
- **Products**: 30 (equity index, energy, metals, grains, softs, rates,
currencies).
- **Tenors**: V0 (front month), V1, V2 for each product.
- **Size on disk**: ~400 MB total (hourly hive partitions + daily
aggregate).
- **Runtime**: ~20-40 minutes for a full refresh (Databento API is fast;
the bottleneck is download volume, not rate limiting).
- **API key**: `DATABENTO_API_KEY` required.
- **Cost**: ~$0.05-0.10 per product per year. A full 30-product × 15-year
refresh runs ~$20-50. **Always** run `--estimate-only` first — new
Databento accounts receive $125 free credit which is enough for the
default ES + NQ + CL demo slice but not for a full fetch.
- **License / attribution**: Databento's standard license permits
personal research and analytics. Redistribution of the raw time
series as a product is prohibited; derived analytics are fine. See
https://databento.com/terms.
## Products
| Group | Symbols (30) |
| ------------- | --------------------------------------------------------- |
| Equity Index | ES, NQ, RTY, YM |
| Energy | CL, NG, RB, HO, BZ |
| Metals | GC, SI, HG, PL |
| Grains | ZC, ZW, ZS, ZM, ZL, ZO |
| Softs | KC, CT, SB, CC, OJ |
| Interest Rates| ZB, ZN, ZF, ZT |
| Currencies | 6E, 6J |
## Download
```bash
# === Market (Databento — paid, always estimate cost first) ===
uv run python data/futures/market/download.py --estimate-only
uv run python data/futures/market/download.py # full
uv run python data/futures/market/download.py --product ES --product NQ
uv run python data/futures/market/download.py --start-date 2020-01-01 --end-date 2023-12-31
# Individual contracts (ES and CL), needed by 02_futures_continuous to
# reconstruct the roll. download.py fetches only the continuous series, which
# have already been rolled, so the roll cannot be taught from them.
uv run python data/futures/market/databento_individual.py --estimate
uv run python data/futures/market/databento_individual.py # ES and CL
uv run python data/futures/market/databento_individual.py --product ES
# === Positioning (CFTC CoT — free, weekly) ===
uv run python data/futures/positioning/cot_download.py # all products, 2020-current
uv run python data/futures/positioning/cot_download.py --products ES,NQ,CL,GC --start-year 2010
```
Output layout under `$ML4T_DATA_PATH/futures/`:
```
market/
├── continuous/
│ ├── hourly/product=<PROD>/year=<YYYY>/data.parquet # raw from Databento
│ └── daily/continuous_daily.parquet # session-aligned daily
├── individual/{PRODUCT}/data.parquet # individual contract roll demo
└── config.yaml # product list, tenors, Databento codes
positioning/
└── cot/{PRODUCT}.parquet # CFTC Commitment of Traders
```
## CFTC Commitment of Traders (free)
Weekly positioning snapshots (Tuesday; released Friday) broken down by
trader category. Used in Ch4 NB 10 for sentiment/positioning features.
```python
from data.futures.loader import load_cot
df = load_cot(products=["ES"], start_date="2020-01-01", end_date="2024-12-31")
df = load_cot() # everything available locally
```
Schema includes `product`, `report_type`, `report_date`, `open_interest`,
and per-trader long/short/net columns (financial: `dealer_*`, `asset_mgr_*`,
`lev_money_*`; commodity: `commercial_*`, `managed_money_*`, `swap_*`).
## Loading
```python
from data import load_cme_futures
df = load_cme_futures() # daily, all products, front + 2 tenors
df = load_cme_futures(frequency="hourly") # hourly panel
df = load_cme_futures(products=["ES", "NQ", "CL"])
df = load_cme_futures(tenors=[0]) # front month only
```
Schema (canonical — note `product` instead of `symbol` for CME per the
book's naming convention):
| Column | Type | Description |
| ----------- | -------- | ------------------------------------- |
| `product` | String | CME product code (e.g., ES) |
| `tenor` | Int | 0 = front month, 1 = next, 2 = after |
| `timestamp` | Datetime | Bar timestamp (daily or hourly) |
| `open` | Float | Opening price |
| `high` | Float | High price |
| `low` | Float | Low price |
| `close` | Float | Closing price |
| `volume` | Int | Trading volume |
## Consumers
- **Ch2**: `05_futures_session_aggregation.py`, `06_cme_futures_eda.py`.
- **Ch6**: `03_cme_futures_setup.py` (carry strategy definition).
- **Ch12-17**: modelling and backtesting.
- **`case_studies/cme_futures/`**: full pipeline from `01_feasibility_analysis.py`
through `17_strategy_analysis.py`.在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。