다중 자산 ETF 후보군 구성 및 필터링
노트북 Machine Learning for Trading
요약
이 문서는 모멘텀 및 자산 간 연구를 위한 후보군으로 쓰이는 일별 ETF 데이터셋을 설명합니다. US와 해외 주식, 채권, 원자재, 특수 펀드, 통화 등 9개 범주를 다룹니다. 데이터는 Yahoo Finance에서 가져오며, 함께 제공되는 워크플로는 데이터를 내려받고 저장하고 불러오고 필터링한 뒤 프로파일링합니다. 종목 및 범주별 커버리지를 확인해 이력과 거래량의 차이를 파악할 수 있습니다.
이 투자 종목군은 바로 거래할 포트폴리오로 제시되지 않습니다. 이후 전략 정의 단계에서 유동성, 이용 가능한 이력, 상관관계 군집을 기준으로 상품을 선택할 예정입니다. 수정 종가는 배당과 액면분할을 반영하며 거래량은 ETF의 거래량을 나타냅니다. 일부 펀드는 이력이 더 짧습니다. 이 문서는 데이터 갱신을 위한 운영 지침을 제공하지만, 수익률 분석이나 거래 결과 또는 이 후보군이 특정 전략에 적합하다는 증거는 제시하지 않습니다.
핵심 아이디어
- ETF 투자 종목군에는 주식, 채권, 원자재, 통화 범주가 포함됩니다.
- 전략을 선택하기 전에 종목별 데이터 커버리지와 평균 거래량을 확인할 수 있습니다.
- 후보군은 유동성, 이력, 상관관계를 기준으로 필터링하도록 설계되었습니다.
- 수정 종가는 배당과 액면분할을 반영하며 일부 ETF는 기록 기간이 짧습니다.
태그
전문
# ETF Universe Dataset
# ETF Universe Dataset
100 diversified ETFs across 9 categories for momentum and cross-asset strategies.
| Property | Value |
|----------|-------|
| **Provider** | Yahoo Finance |
| **Asset Class** | Multi-asset (Equity, Fixed Income, Commodities, Currency) |
| **Frequency** | Daily |
| **Symbols** | 100 ETFs |
| **Coverage** | 2006-2025 |
| **Size** | ~16 MB |
| **API Key** | None (free) |
| **Loader** | `load_etfs()` |
```python
"""ETF Universe - download, explore, and update workflow."""
import json
from pathlib import Path
import polars as pl
import yaml
```
## 1. Configuration
The ETF universe is defined in `config.yaml`. This is the **candidate pool** -
strategy definition (Chapter 6) filters this down based on liquidity, history,
and correlation clustering.
```python
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]
print("=== ETF Configuration ===")
print(f"Provider: {etf_config['provider']}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Frequency: {etf_config['frequency']}")
print(f"\nCategories ({len(etf_config['tickers'])}):")
for category, info in etf_config["tickers"].items():
symbols = info["symbols"]
print(f" {category}: {len(symbols)} ETFs")
total_etfs = sum(len(info["symbols"]) for info in etf_config["tickers"].values())
print(f"\nTotal: {total_etfs} ETFs")
```
## 2. API Key Setup
**No API key required.** Yahoo Finance data is free and publicly accessible.
The `ml4t-data` library handles rate limiting automatically to avoid
being blocked by Yahoo Finance.
```python
print("Yahoo Finance requires no API key - data is publicly available.")
```
## 3. Download Data
The download uses the `ml4t-data` library which handles:
- Rate limiting (1 second delay between batches)
- Retry logic for failed requests
- Consistent schema output
**Note**: First-time download takes ~2-3 minutes for 100 ETFs.
```python
def download_etf_data(dry_run: bool = False, force: bool = False, symbols: list[str] | None = None):
"""Download ETF data from Yahoo Finance.
Args:
dry_run: If True, show what would be downloaded without doing it
force: If True, re-download even if data exists
symbols: Specific symbols to download (default: all from config)
"""
from ml4t.data.providers import YahooFinanceProvider
from utils import ML4T_DATA_PATH
# Load config
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]
# Flatten symbols list
if symbols is None:
symbols = []
for category_info in etf_config["tickers"].values():
symbols.extend(category_info["symbols"])
output_dir = ML4T_DATA_PATH / "etfs" / "market"
output_path = output_dir / "etf_universe.parquet"
print("=== ETF Download ===")
print(f"Symbols: {len(symbols)}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Output: {output_path}")
if dry_run:
print("\n[DRY RUN] Would download:")
for i, symbol in enumerate(symbols, 1):
print(f" {i:3}. {symbol}")
return
# Check existing data
if output_path.exists() and not force:
existing = pl.read_parquet(output_path)
existing_symbols = set(existing["symbol"].unique().to_list())
missing = [s for s in symbols if s not in existing_symbols]
if not missing:
print(f"\nAll {len(symbols)} ETFs already downloaded.")
print("Use force=True to re-download.")
return existing
print(f"Found {len(existing_symbols)} existing, downloading {len(missing)} missing...")
symbols = missing
# Initialize provider and download
provider = YahooFinanceProvider()
print(f"\nDownloading {len(symbols)} ETFs...")
etf_data = provider.fetch_batch_ohlcv(
symbols=symbols,
start=etf_config["start"],
end=etf_config["end"],
frequency="daily",
chunk_size=50,
delay_seconds=1.0,
)
# Combine with existing data if applicable
if output_path.exists() and not force:
existing = pl.read_parquet(output_path)
etf_data = pl.concat([existing, etf_data])
# Save
output_dir.mkdir(parents=True, exist_ok=True)
etf_data.write_parquet(output_path)
print("\n=== Complete ===")
print(f"Total rows: {len(etf_data):,}")
print(f"Symbols: {etf_data['symbol'].n_unique()}")
print(f"Date range: {etf_data['timestamp'].min()} to {etf_data['timestamp'].max()}")
print(f"Saved to: {output_path}")
return etf_data
```
### Download All ETFs
```python
# Uncomment to download all ETF data
# download_etf_data()
```
### Dry Run (Preview)
See what would be downloaded without actually downloading:
```python
download_etf_data(dry_run=True)
```
## 4. Load and Explore
Once downloaded, use the loader throughout the book:
```python
from data import load_etfs
# Load all ETF data
df = load_etfs()
print(f"Shape: {df.shape}")
print(f"Symbols: {df['symbol'].n_unique()}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")
```
```python
# Schema
df.schema
```
```python
# Preview
df.head(10)
```
### Coverage by Symbol
```python
# Coverage and basic stats by symbol
coverage = (
df.group_by("symbol")
.agg(
pl.col("timestamp").min().alias("first_date"),
pl.col("timestamp").max().alias("last_date"),
pl.len().alias("n_bars"),
pl.col("volume").mean().alias("avg_daily_volume"),
)
.sort("avg_daily_volume", descending=True)
)
coverage.head(20)
```
### Category Summary
```python
# Build category mapping from config
category_map = {}
for category, info in etf_config["tickers"].items():
for symbol in info["symbols"]:
category_map[symbol] = category
df_with_cat = df.with_columns(pl.col("symbol").replace(category_map).alias("category"))
category_summary = (
df_with_cat.group_by("category")
.agg(
pl.col("symbol").n_unique().alias("n_symbols"),
pl.col("timestamp").min().alias("earliest"),
pl.col("timestamp").max().alias("latest"),
pl.col("volume").mean().alias("avg_volume"),
)
.sort("n_symbols", descending=True)
)
category_summary
```
## 5. Data Profile
Profiles document the dataset structure, statistics, and quality metrics.
They are stored alongside the data files.
```python
from ml4t.data.storage.data_profile import load_profile
from utils import ML4T_DATA_PATH
profile_path = ML4T_DATA_PATH / "etfs" / "market" / "etf_universe_profile.json"
profile = load_profile(profile_path)
if profile is None:
print(f"No profile at {profile_path}")
print(
"Profiles are written next to the data by whatever builds the dataset - the\n"
"download script in this directory, or the ml4t-data loader it drives - through\n"
"ml4t.data.storage.data_profile. There is no separate profile-generating script,\n"
"and nothing in this notebook writes one."
)
else:
print("=== ETF Universe Profile ===")
print(f"Written by {profile.source}")
print(profile.summary())
```
## 6. Loader Options
The loader supports filtering by symbols and date range:
```python
# Specific symbols
spy_qqq = load_etfs(symbols=["SPY", "QQQ"])
print(f"SPY + QQQ only: {spy_qqq.shape}")
```
```python
# Date range
recent = load_etfs(start_date="2024-01-01")
print(f"2024 onwards: {recent.shape}")
```
```python
# Combined filters
filtered = load_etfs(
symbols=["SPY", "QQQ", "IWM", "TLT", "GLD"], start_date="2020-01-01", end_date="2023-12-31"
)
print(f"5 ETFs, 2020-2023: {filtered.shape}")
```
## 7. Documentation
### Yahoo Finance
- [Yahoo Finance API (unofficial)](https://python-yahoofinance.readthedocs.io/)
### ETF Categories
| Category | Count | Description |
|----------|-------|-------------|
| `us_equity_broad` | 10 | Large, mid, small cap, equal weight |
| `us_equity_style` | 10 | Value, growth, momentum, dividend |
| `us_sectors` | 13 | SPDR sector ETFs + real estate |
| `international_developed` | 18 | EAFE, Europe, Japan, country ETFs |
| `emerging_markets` | 11 | EM broad + China, Brazil, India, etc. |
| `fixed_income` | 15 | Treasury, corporate, high yield, TIPS |
| `commodities` | 9 | Gold, silver, oil, broad commodity |
| `specialty` | 10 | Biotech, semiconductors, regional banks |
| `currency` | 4 | USD, EUR, JPY, GBP currency ETFs |
### Data Quality Notes
- Volume represents actual ETF trading volume
- Adjusted close accounts for dividends and splits
- Some ETFs have shorter history (check `first_date` in coverage)
## 8. Updating Data
To update with the latest data, re-run the download:
```python
# Update to latest available data
download_etf_data()
# Force full re-download
download_etf_data(force=True)
```
**Tip**: Update the `end` date in `config.yaml` before re-downloading
to extend the coverage period.
## Summary
| Item | Value |
|------|-------|
| Symbols | 100 ETFs across 9 categories |
| Frequency | Daily |
| Coverage | 2006-2025 |
| Provider | Yahoo Finance (free) |
| Config | `config.yaml` |
| Loader | `load_etfs(symbols, start_date, end_date)` |
| Profile | `$ML4T_DATA_PATH/etfs/market/etf_universe_profile.json` |
**Note**: This is the **candidate pool**. Chapter 6 filters to ~80 ETFs
based on liquidity, history, and correlation clustering.출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.