跳至正文
返回文库全部文档

构建并筛选多资产ETF候选标的池

笔记本 《交易机器学习》

总结

本文介绍一个每日ETF数据集,作为动量和跨资产研究的候选标的池。它涵盖九大类别,包括US和国际股票、固定收益、商品、特色基金及货币。数据来自Yahoo Finance,配套流程会下载、存储、加载、筛选并概览记录。您可以按代码和类别检查数据覆盖情况,以了解历史记录和交易量的差异。

本文没有将该标的池描述为可直接交易的投资组合。后续策略定义预计会根据流动性、历史数据可用性和相关性聚类来筛选工具。复权收盘价包含股息和拆股调整,成交量反映ETF交易量;部分基金的历史记录较短。本文提供数据更新的运维指引,但不包含收益分析或交易结果,也没有证据表明这组候选标的适用于任何特定策略。

核心观点

  • ETF标的池涵盖股票、固定收益、商品和货币类别。
  • 选择策略前,可以按代码检查数据覆盖情况和平均成交量。
  • 该候选池旨在根据流动性、历史数据和相关性进行筛选。
  • 复权收盘价已考虑股息和拆股,但部分ETF的历史记录较短。

标签

全文
# ETF Universe Dataset


# ETF Universe Dataset

100 diversified ETFs across 9 categories for momentum and cross-asset strategies.

| Property | Value |
|----------|-------|
| **Provider** | Yahoo Finance |
| **Asset Class** | Multi-asset (Equity, Fixed Income, Commodities, Currency) |
| **Frequency** | Daily |
| **Symbols** | 100 ETFs |
| **Coverage** | 2006-2025 |
| **Size** | ~16 MB |
| **API Key** | None (free) |
| **Loader** | `load_etfs()` |

```python
"""ETF Universe - download, explore, and update workflow."""

import json
from pathlib import Path

import polars as pl
import yaml
```

## 1. Configuration

The ETF universe is defined in `config.yaml`. This is the **candidate pool** -
strategy definition (Chapter 6) filters this down based on liquidity, history,
and correlation clustering.

```python
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]

print("=== ETF Configuration ===")
print(f"Provider: {etf_config['provider']}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Frequency: {etf_config['frequency']}")
print(f"\nCategories ({len(etf_config['tickers'])}):")
for category, info in etf_config["tickers"].items():
    symbols = info["symbols"]
    print(f"  {category}: {len(symbols)} ETFs")

total_etfs = sum(len(info["symbols"]) for info in etf_config["tickers"].values())
print(f"\nTotal: {total_etfs} ETFs")
```

## 2. API Key Setup

**No API key required.** Yahoo Finance data is free and publicly accessible.

The `ml4t-data` library handles rate limiting automatically to avoid
being blocked by Yahoo Finance.

```python
print("Yahoo Finance requires no API key - data is publicly available.")
```

## 3. Download Data

The download uses the `ml4t-data` library which handles:
- Rate limiting (1 second delay between batches)
- Retry logic for failed requests
- Consistent schema output

**Note**: First-time download takes ~2-3 minutes for 100 ETFs.

```python
def download_etf_data(dry_run: bool = False, force: bool = False, symbols: list[str] | None = None):
    """Download ETF data from Yahoo Finance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from ml4t.data.providers import YahooFinanceProvider

    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    etf_config = config["etfs"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in etf_config["tickers"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "etfs" / "market"
    output_path = output_dir / "etf_universe.parquet"

    print("=== ETF Download ===")
    print(f"Symbols: {len(symbols)}")
    print(f"Date range: {etf_config['start']} to {etf_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for i, symbol in enumerate(symbols, 1):
            print(f"  {i:3}. {symbol}")
        return

    # Check existing data
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        existing_symbols = set(existing["symbol"].unique().to_list())
        missing = [s for s in symbols if s not in existing_symbols]
        if not missing:
            print(f"\nAll {len(symbols)} ETFs already downloaded.")
            print("Use force=True to re-download.")
            return existing
        print(f"Found {len(existing_symbols)} existing, downloading {len(missing)} missing...")
        symbols = missing

    # Initialize provider and download
    provider = YahooFinanceProvider()
    print(f"\nDownloading {len(symbols)} ETFs...")

    etf_data = provider.fetch_batch_ohlcv(
        symbols=symbols,
        start=etf_config["start"],
        end=etf_config["end"],
        frequency="daily",
        chunk_size=50,
        delay_seconds=1.0,
    )

    # Combine with existing data if applicable
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        etf_data = pl.concat([existing, etf_data])

    # Save
    output_dir.mkdir(parents=True, exist_ok=True)
    etf_data.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(etf_data):,}")
    print(f"Symbols: {etf_data['symbol'].n_unique()}")
    print(f"Date range: {etf_data['timestamp'].min()} to {etf_data['timestamp'].max()}")
    print(f"Saved to: {output_path}")

    return etf_data
```

### Download All ETFs

```python
# Uncomment to download all ETF data
# download_etf_data()
```

### Dry Run (Preview)

See what would be downloaded without actually downloading:

```python
download_etf_data(dry_run=True)
```

## 4. Load and Explore

Once downloaded, use the loader throughout the book:

```python
from data import load_etfs

# Load all ETF data
df = load_etfs()

print(f"Shape: {df.shape}")
print(f"Symbols: {df['symbol'].n_unique()}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")
```

```python
# Schema
df.schema
```

```python
# Preview
df.head(10)
```

### Coverage by Symbol

```python
# Coverage and basic stats by symbol
coverage = (
    df.group_by("symbol")
    .agg(
        pl.col("timestamp").min().alias("first_date"),
        pl.col("timestamp").max().alias("last_date"),
        pl.len().alias("n_bars"),
        pl.col("volume").mean().alias("avg_daily_volume"),
    )
    .sort("avg_daily_volume", descending=True)
)
coverage.head(20)
```

### Category Summary

```python
# Build category mapping from config
category_map = {}
for category, info in etf_config["tickers"].items():
    for symbol in info["symbols"]:
        category_map[symbol] = category

df_with_cat = df.with_columns(pl.col("symbol").replace(category_map).alias("category"))

category_summary = (
    df_with_cat.group_by("category")
    .agg(
        pl.col("symbol").n_unique().alias("n_symbols"),
        pl.col("timestamp").min().alias("earliest"),
        pl.col("timestamp").max().alias("latest"),
        pl.col("volume").mean().alias("avg_volume"),
    )
    .sort("n_symbols", descending=True)
)
category_summary
```

## 5. Data Profile

Profiles document the dataset structure, statistics, and quality metrics.
They are stored alongside the data files.

```python
from ml4t.data.storage.data_profile import load_profile

from utils import ML4T_DATA_PATH

profile_path = ML4T_DATA_PATH / "etfs" / "market" / "etf_universe_profile.json"
profile = load_profile(profile_path)

if profile is None:
    print(f"No profile at {profile_path}")
    print(
        "Profiles are written next to the data by whatever builds the dataset - the\n"
        "download script in this directory, or the ml4t-data loader it drives - through\n"
        "ml4t.data.storage.data_profile. There is no separate profile-generating script,\n"
        "and nothing in this notebook writes one."
    )
else:
    print("=== ETF Universe Profile ===")
    print(f"Written by {profile.source}")
    print(profile.summary())
```

## 6. Loader Options

The loader supports filtering by symbols and date range:

```python
# Specific symbols
spy_qqq = load_etfs(symbols=["SPY", "QQQ"])
print(f"SPY + QQQ only: {spy_qqq.shape}")
```

```python
# Date range
recent = load_etfs(start_date="2024-01-01")
print(f"2024 onwards: {recent.shape}")
```

```python
# Combined filters
filtered = load_etfs(
    symbols=["SPY", "QQQ", "IWM", "TLT", "GLD"], start_date="2020-01-01", end_date="2023-12-31"
)
print(f"5 ETFs, 2020-2023: {filtered.shape}")
```

## 7. Documentation

### Yahoo Finance
- [Yahoo Finance API (unofficial)](https://python-yahoofinance.readthedocs.io/)


### ETF Categories

| Category | Count | Description |
|----------|-------|-------------|
| `us_equity_broad` | 10 | Large, mid, small cap, equal weight |
| `us_equity_style` | 10 | Value, growth, momentum, dividend |
| `us_sectors` | 13 | SPDR sector ETFs + real estate |
| `international_developed` | 18 | EAFE, Europe, Japan, country ETFs |
| `emerging_markets` | 11 | EM broad + China, Brazil, India, etc. |
| `fixed_income` | 15 | Treasury, corporate, high yield, TIPS |
| `commodities` | 9 | Gold, silver, oil, broad commodity |
| `specialty` | 10 | Biotech, semiconductors, regional banks |
| `currency` | 4 | USD, EUR, JPY, GBP currency ETFs |

### Data Quality Notes
- Volume represents actual ETF trading volume
- Adjusted close accounts for dividends and splits
- Some ETFs have shorter history (check `first_date` in coverage)

## 8. Updating Data

To update with the latest data, re-run the download:

```python
# Update to latest available data
download_etf_data()

# Force full re-download
download_etf_data(force=True)
```

**Tip**: Update the `end` date in `config.yaml` before re-downloading
to extend the coverage period.

## Summary

| Item | Value |
|------|-------|
| Symbols | 100 ETFs across 9 categories |
| Frequency | Daily |
| Coverage | 2006-2025 |
| Provider | Yahoo Finance (free) |
| Config | `config.yaml` |
| Loader | `load_etfs(symbols, start_date, end_date)` |
| Profile | `$ML4T_DATA_PATH/etfs/market/etf_universe_profile.json` |

**Note**: This is the **candidate pool**. Chapter 6 filters to ~80 ETFs
based on liquidity, history, and correlation clustering.

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。