선물 분석용 CFTC 포지션 데이터 다운로드
코드 Machine Learning for Trading
요약
이 문서는 선물 상품을 선택해 주간 CFTC 트레이더 포지션 데이터인 Commitment of Traders를 가져오고, 상품별 이력을 Parquet 파일로 저장하는 방법을 설명합니다. COT 보고서는 화요일 포지션을 담아 금요일에 공개되며, 트레이더 범주는 금융 선물과 상품 선물에서 다릅니다. 생성된 데이터에는 보고서 날짜, 미결제 약정, 롱·숏·순포지션이 포함되며 선물 조사 과정에서 포지션 또는 심리 특성으로 활용할 수 있습니다.
다운로더에서 상품과 연도 범위를 선택하거나 설정된 상품 목록과 기본 날짜 범위를 사용할 수 있습니다. 요청한 상품 코드를 검증하고, 데이터가 반환되지 않은 상품을 기록하며, 가져오기 실패를 보고합니다. 이 내용은 COT 신호 성과 분석이 아니라 운영 안내입니다. 백테스트, 예측 증거, 포지션 변화를 해석하는 규칙은 제공하지 않습니다. COT 관측값은 주간 스냅샷이므로 특정 매매 전략에 대한 시의성이나 유용성을 입증하지 않습니다.
핵심 아이디어
- CFTC 보고서는 주간 선물 포지션 스냅샷을 제공하며 트레이더 범주는 시장과 보고서 유형에 따라 다릅니다.
- 보고서 날짜별 미결제 약정과 트레이더 범주별 롱·숏·순포지션을 기록합니다.
- 사용자는 상품과 연도를 선택할 수 있으며, 가져오기 실패와 빈 결과를 별도로 보고합니다.
- 이 문서는 데이터 조회와 저장을 설명하지만 COT 특성이 수익률을 예측하는지는 시험하지 않습니다.
태그
전문
# cot_download.py
```py
#!/usr/bin/env python3
"""Download CFTC Commitment of Traders (COT) data.
CFTC publishes weekly COT reports (Tuesday snapshot, released Friday) showing
futures positioning broken down by trader type (dealers, asset managers,
leveraged money for financial futures; commercials, managed money for
commodities). The data is free and useful for sentiment/positioning features.
This downloader uses ``ml4t.data.cot.COTFetcher`` — which wraps the
``cot_reports`` library — to fetch per-product panels and writes one parquet
per product to ``$ML4T_DATA_PATH/futures/positioning/cot/{product}.parquet``. The
``load_cot()`` loader in ``data/futures/loader.py`` consumes these files.
Output layout under ``$ML4T_DATA_PATH/futures/positioning/cot/``::
{PRODUCT}.parquet one parquet per product code (e.g., ES.parquet)
Schema (columns vary by report type but include):
product exchange product code (ES, CL, GC, …)
report_type CFTC report that produced the row
report_date Tuesday snapshot date
open_interest total open interest
<trader>_long long positions per trader category
<trader>_short short positions per trader category
<trader>_net computed long − short per category
Usage::
# Default: all products in PRODUCT_MAPPINGS, 2020–current year
python data/futures/positioning/cot_download.py
# Restrict to a subset
python data/futures/positioning/cot_download.py --products ES,NQ,CL,GC
# Wider year range
python data/futures/positioning/cot_download.py --start-year 2010 --end-year 2024
# Override output root
python data/futures/positioning/cot_download.py --data-path /tmp/ml4t-data
"""
from __future__ import annotations
import argparse
from pathlib import Path
from ml4t.data.cot import PRODUCT_MAPPINGS, COTConfig, COTFetcher
from utils.downloading import resolve_data_dir
def main() -> int:
parser = argparse.ArgumentParser(
description="Download CFTC Commitment of Traders data",
formatter_class=argparse.RawDescriptionHelpFormatter,
)
parser.add_argument(
"--products",
type=str,
default=None,
help=(
"Comma-separated product codes (default: all in PRODUCT_MAPPINGS). "
f"Available: {', '.join(sorted(PRODUCT_MAPPINGS.keys()))}"
),
)
parser.add_argument(
"--start-year",
type=int,
default=2020,
help="First calendar year to fetch (default: 2020)",
)
parser.add_argument(
"--end-year",
type=int,
default=None,
help="Last calendar year to fetch (default: current year)",
)
parser.add_argument(
"--data-path",
type=Path,
default=None,
help="Override output root (default: $ML4T_DATA_PATH)",
)
args = parser.parse_args()
if args.products:
products = [p.strip().upper() for p in args.products.split(",") if p.strip()]
unknown = [p for p in products if p not in PRODUCT_MAPPINGS]
if unknown:
print(f"ERROR: unknown product code(s): {', '.join(unknown)}")
print(f"Available: {', '.join(sorted(PRODUCT_MAPPINGS.keys()))}")
return 1
else:
products = sorted(PRODUCT_MAPPINGS.keys())
data_path = resolve_data_dir(args.data_path)
output_dir = data_path / "futures" / "positioning" / "cot"
output_dir.mkdir(parents=True, exist_ok=True)
config = COTConfig(
products=products,
start_year=args.start_year,
end_year=args.end_year,
storage_path=output_dir,
)
fetcher = COTFetcher(config)
print()
print(f"Output: {output_dir}")
print(f"Years: {config.start_year}–{config.end_year}")
print(f"Products: {len(products)} ({', '.join(products)})")
print()
written = 0
empty = 0
failed: list[str] = []
for i, product in enumerate(products, 1):
print(f" [{i}/{len(products)}] {product}…", end="", flush=True)
try:
df = fetcher.fetch_product(product)
except Exception as e:
failed.append(product)
print(f" FAILED ({e})")
continue
if df.is_empty():
empty += 1
print(" no rows returned")
continue
out_path = output_dir / f"{product}.parquet"
df.write_parquet(out_path)
written += 1
print(f" {len(df):,} rows → {out_path.name}")
print()
print(f"Wrote: {written} parquet(s)")
if empty:
print(f"Empty: {empty} product(s) returned no rows")
if failed:
print(f"Failed: {len(failed)} — {', '.join(failed)}")
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.