기간 구조 및 캐리 연구를 위한 CME 선물 데이터
기사 Machine Learning for Trading
요약
시간봉과 일봉, 연속 근월물 계약, 여러 상품군의 이연 만기 두 구간을 포함한 CME 선물 데이터셋을 설명합니다. 데이터셋의 범위, 필드, 불러오기 옵션, 관련 주간 CFTC 포지션 데이터도 다룹니다. 만기 구조는 기간 구조 분석과 캐리 연구를 지원하며, 롤을 재구성하고 연구하려면 개별 계약 기록이 필요합니다.
자료는 트레이딩 연구가 아니라 주로 데이터셋 안내입니다. 전략 결과나 특정 접근법의 수익성 증거는 제공하지 않습니다. 시장 이력을 사용하려면 유료 데이터 소스와 API 키가 필요하며 전체 새로고침에는 상당한 비용이 발생할 수 있습니다. 연속 계약은 이미 롤 처리되어 있으므로 그 데이터만으로는 롤 구성 방식을 확인할 수 없습니다. CFTC 데이터는 트레이더 범주별 포지션을 추가로 제공하지만 주간 단위이며 발표가 늦습니다. 원시 데이터의 재배포는 소스 라이선스에 따라 제한됩니다.
핵심 아이디어
- 데이터셋은 근월물과 이연 만기 두 구간의 시간별·일별 CME 선물 봉을 제공합니다.
- 만기 비교는 기간 구조 및 캐리 분석에 활용할 수 있습니다.
- 연속 시계열은 이미 롤 처리되어 있으므로 롤을 재구성하려면 개별 계약 데이터가 필요합니다.
- 주간 CFTC 보고서는 발표 지연이 있는 트레이더 범주별 포지션 스냅샷을 제공합니다.
- 시장 데이터에는 유료 API 소스가 필요하며, 원시 시계열의 재배포는 소스에서 제한합니다.
태그
전문
# CME Futures (Databento)
# CME Futures (Databento)
30 CME futures products — continuous front-month contracts plus the next
two tenors — for term-structure analysis, carry strategies, and the
`cme_futures` case study. Daily and hourly bars available.
## Dataset
- **Source**: Databento via the `GLBX.MDP3` CME dataset.
- **Coverage**: 2011-01-01 → present, hourly OHLCV with daily aggregation.
- **Products**: 30 (equity index, energy, metals, grains, softs, rates,
currencies).
- **Tenors**: V0 (front month), V1, V2 for each product.
- **Size on disk**: ~400 MB total (hourly hive partitions + daily
aggregate).
- **Runtime**: ~20-40 minutes for a full refresh (Databento API is fast;
the bottleneck is download volume, not rate limiting).
- **API key**: `DATABENTO_API_KEY` required.
- **Cost**: ~$0.05-0.10 per product per year. A full 30-product × 15-year
refresh runs ~$20-50. **Always** run `--estimate-only` first — new
Databento accounts receive $125 free credit which is enough for the
default ES + NQ + CL demo slice but not for a full fetch.
- **License / attribution**: Databento's standard license permits
personal research and analytics. Redistribution of the raw time
series as a product is prohibited; derived analytics are fine. See
https://databento.com/terms.
## Products
| Group | Symbols (30) |
| ------------- | --------------------------------------------------------- |
| Equity Index | ES, NQ, RTY, YM |
| Energy | CL, NG, RB, HO, BZ |
| Metals | GC, SI, HG, PL |
| Grains | ZC, ZW, ZS, ZM, ZL, ZO |
| Softs | KC, CT, SB, CC, OJ |
| Interest Rates| ZB, ZN, ZF, ZT |
| Currencies | 6E, 6J |
## Download
```bash
# === Market (Databento — paid, always estimate cost first) ===
uv run python data/futures/market/download.py --estimate-only
uv run python data/futures/market/download.py # full
uv run python data/futures/market/download.py --product ES --product NQ
uv run python data/futures/market/download.py --start-date 2020-01-01 --end-date 2023-12-31
# Individual contracts (ES and CL), needed by 02_futures_continuous to
# reconstruct the roll. download.py fetches only the continuous series, which
# have already been rolled, so the roll cannot be taught from them.
uv run python data/futures/market/databento_individual.py --estimate
uv run python data/futures/market/databento_individual.py # ES and CL
uv run python data/futures/market/databento_individual.py --product ES
# === Positioning (CFTC CoT — free, weekly) ===
uv run python data/futures/positioning/cot_download.py # all products, 2020-current
uv run python data/futures/positioning/cot_download.py --products ES,NQ,CL,GC --start-year 2010
```
Output layout under `$ML4T_DATA_PATH/futures/`:
```
market/
├── continuous/
│ ├── hourly/product=<PROD>/year=<YYYY>/data.parquet # raw from Databento
│ └── daily/continuous_daily.parquet # session-aligned daily
├── individual/{PRODUCT}/data.parquet # individual contract roll demo
└── config.yaml # product list, tenors, Databento codes
positioning/
└── cot/{PRODUCT}.parquet # CFTC Commitment of Traders
```
## CFTC Commitment of Traders (free)
Weekly positioning snapshots (Tuesday; released Friday) broken down by
trader category. Used in Ch4 NB 10 for sentiment/positioning features.
```python
from data.futures.loader import load_cot
df = load_cot(products=["ES"], start_date="2020-01-01", end_date="2024-12-31")
df = load_cot() # everything available locally
```
Schema includes `product`, `report_type`, `report_date`, `open_interest`,
and per-trader long/short/net columns (financial: `dealer_*`, `asset_mgr_*`,
`lev_money_*`; commodity: `commercial_*`, `managed_money_*`, `swap_*`).
## Loading
```python
from data import load_cme_futures
df = load_cme_futures() # daily, all products, front + 2 tenors
df = load_cme_futures(frequency="hourly") # hourly panel
df = load_cme_futures(products=["ES", "NQ", "CL"])
df = load_cme_futures(tenors=[0]) # front month only
```
Schema (canonical — note `product` instead of `symbol` for CME per the
book's naming convention):
| Column | Type | Description |
| ----------- | -------- | ------------------------------------- |
| `product` | String | CME product code (e.g., ES) |
| `tenor` | Int | 0 = front month, 1 = next, 2 = after |
| `timestamp` | Datetime | Bar timestamp (daily or hourly) |
| `open` | Float | Opening price |
| `high` | Float | High price |
| `low` | Float | Low price |
| `close` | Float | Closing price |
| `volume` | Int | Trading volume |
## Consumers
- **Ch2**: `05_futures_session_aggregation.py`, `06_cme_futures_eda.py`.
- **Ch6**: `03_cme_futures_setup.py` (carry strategy definition).
- **Ch12-17**: modelling and backtesting.
- **`case_studies/cme_futures/`**: full pipeline from `01_feasibility_analysis.py`
through `17_strategy_analysis.py`.출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.