Chuyển đến nội dung
Tất cả tài liệu trong thư viện

Xây dựng và lọc tập ứng viên đa tài sản ETF

Notebook Machine Learning for Trading

Tóm tắt

Tài liệu này mô tả tập dữ liệu ETF hàng ngày, được dùng làm nhóm ứng viên cho nghiên cứu động lượng và đa tài sản. Tập dữ liệu bao gồm chín nhóm, trong đó có cổ phiếu US và quốc tế, thu nhập cố định, hàng hóa, quỹ chuyên biệt và tiền tệ. Dữ liệu lấy từ Yahoo Finance; quy trình đi kèm tải xuống, lưu trữ, nạp, lọc và lập hồ sơ các bản ghi. Có thể kiểm tra độ bao phủ theo mã và nhóm để nhận diện lịch sử và khối lượng giao dịch khác nhau.

Tập tài sản này không được giới thiệu như một danh mục sẵn sàng giao dịch. Định nghĩa chiến lược sau này dự kiến sẽ chọn công cụ dựa trên thanh khoản, lịch sử khả dụng và phân cụm tương quan. Giá đóng cửa điều chỉnh đã tính cổ tức và chia tách, còn khối lượng phản ánh khối lượng giao dịch ETF; một số quỹ có lịch sử ngắn hơn. Tài liệu cung cấp hướng dẫn vận hành để cập nhật dữ liệu nhưng không có phân tích lợi nhuận, kết quả giao dịch hay bằng chứng rằng nhóm ứng viên này phù hợp với một chiến lược cụ thể nào.

Ý chính

  • Tập tài sản ETF bao gồm các nhóm cổ phiếu, thu nhập cố định, hàng hóa và tiền tệ.
  • Có thể xem độ bao phủ dữ liệu và khối lượng trung bình theo mã trước khi chọn chiến lược.
  • Nhóm ứng viên được thiết kế để lọc theo thanh khoản, lịch sử và tương quan.
  • Giá đóng cửa điều chỉnh tính đến cổ tức và chia tách, trong khi một số ETF có dữ liệu lịch sử ngắn hơn.

Thẻ

Toàn văn
# ETF Universe Dataset


# ETF Universe Dataset

100 diversified ETFs across 9 categories for momentum and cross-asset strategies.

| Property | Value |
|----------|-------|
| **Provider** | Yahoo Finance |
| **Asset Class** | Multi-asset (Equity, Fixed Income, Commodities, Currency) |
| **Frequency** | Daily |
| **Symbols** | 100 ETFs |
| **Coverage** | 2006-2025 |
| **Size** | ~16 MB |
| **API Key** | None (free) |
| **Loader** | `load_etfs()` |

```python
"""ETF Universe - download, explore, and update workflow."""

import json
from pathlib import Path

import polars as pl
import yaml
```

## 1. Configuration

The ETF universe is defined in `config.yaml`. This is the **candidate pool** -
strategy definition (Chapter 6) filters this down based on liquidity, history,
and correlation clustering.

```python
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]

print("=== ETF Configuration ===")
print(f"Provider: {etf_config['provider']}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Frequency: {etf_config['frequency']}")
print(f"\nCategories ({len(etf_config['tickers'])}):")
for category, info in etf_config["tickers"].items():
    symbols = info["symbols"]
    print(f"  {category}: {len(symbols)} ETFs")

total_etfs = sum(len(info["symbols"]) for info in etf_config["tickers"].values())
print(f"\nTotal: {total_etfs} ETFs")
```

## 2. API Key Setup

**No API key required.** Yahoo Finance data is free and publicly accessible.

The `ml4t-data` library handles rate limiting automatically to avoid
being blocked by Yahoo Finance.

```python
print("Yahoo Finance requires no API key - data is publicly available.")
```

## 3. Download Data

The download uses the `ml4t-data` library which handles:
- Rate limiting (1 second delay between batches)
- Retry logic for failed requests
- Consistent schema output

**Note**: First-time download takes ~2-3 minutes for 100 ETFs.

```python
def download_etf_data(dry_run: bool = False, force: bool = False, symbols: list[str] | None = None):
    """Download ETF data from Yahoo Finance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from ml4t.data.providers import YahooFinanceProvider

    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    etf_config = config["etfs"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in etf_config["tickers"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "etfs" / "market"
    output_path = output_dir / "etf_universe.parquet"

    print("=== ETF Download ===")
    print(f"Symbols: {len(symbols)}")
    print(f"Date range: {etf_config['start']} to {etf_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for i, symbol in enumerate(symbols, 1):
            print(f"  {i:3}. {symbol}")
        return

    # Check existing data
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        existing_symbols = set(existing["symbol"].unique().to_list())
        missing = [s for s in symbols if s not in existing_symbols]
        if not missing:
            print(f"\nAll {len(symbols)} ETFs already downloaded.")
            print("Use force=True to re-download.")
            return existing
        print(f"Found {len(existing_symbols)} existing, downloading {len(missing)} missing...")
        symbols = missing

    # Initialize provider and download
    provider = YahooFinanceProvider()
    print(f"\nDownloading {len(symbols)} ETFs...")

    etf_data = provider.fetch_batch_ohlcv(
        symbols=symbols,
        start=etf_config["start"],
        end=etf_config["end"],
        frequency="daily",
        chunk_size=50,
        delay_seconds=1.0,
    )

    # Combine with existing data if applicable
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        etf_data = pl.concat([existing, etf_data])

    # Save
    output_dir.mkdir(parents=True, exist_ok=True)
    etf_data.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(etf_data):,}")
    print(f"Symbols: {etf_data['symbol'].n_unique()}")
    print(f"Date range: {etf_data['timestamp'].min()} to {etf_data['timestamp'].max()}")
    print(f"Saved to: {output_path}")

    return etf_data
```

### Download All ETFs

```python
# Uncomment to download all ETF data
# download_etf_data()
```

### Dry Run (Preview)

See what would be downloaded without actually downloading:

```python
download_etf_data(dry_run=True)
```

## 4. Load and Explore

Once downloaded, use the loader throughout the book:

```python
from data import load_etfs

# Load all ETF data
df = load_etfs()

print(f"Shape: {df.shape}")
print(f"Symbols: {df['symbol'].n_unique()}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")
```

```python
# Schema
df.schema
```

```python
# Preview
df.head(10)
```

### Coverage by Symbol

```python
# Coverage and basic stats by symbol
coverage = (
    df.group_by("symbol")
    .agg(
        pl.col("timestamp").min().alias("first_date"),
        pl.col("timestamp").max().alias("last_date"),
        pl.len().alias("n_bars"),
        pl.col("volume").mean().alias("avg_daily_volume"),
    )
    .sort("avg_daily_volume", descending=True)
)
coverage.head(20)
```

### Category Summary

```python
# Build category mapping from config
category_map = {}
for category, info in etf_config["tickers"].items():
    for symbol in info["symbols"]:
        category_map[symbol] = category

df_with_cat = df.with_columns(pl.col("symbol").replace(category_map).alias("category"))

category_summary = (
    df_with_cat.group_by("category")
    .agg(
        pl.col("symbol").n_unique().alias("n_symbols"),
        pl.col("timestamp").min().alias("earliest"),
        pl.col("timestamp").max().alias("latest"),
        pl.col("volume").mean().alias("avg_volume"),
    )
    .sort("n_symbols", descending=True)
)
category_summary
```

## 5. Data Profile

Profiles document the dataset structure, statistics, and quality metrics.
They are stored alongside the data files.

```python
from ml4t.data.storage.data_profile import load_profile

from utils import ML4T_DATA_PATH

profile_path = ML4T_DATA_PATH / "etfs" / "market" / "etf_universe_profile.json"
profile = load_profile(profile_path)

if profile is None:
    print(f"No profile at {profile_path}")
    print(
        "Profiles are written next to the data by whatever builds the dataset - the\n"
        "download script in this directory, or the ml4t-data loader it drives - through\n"
        "ml4t.data.storage.data_profile. There is no separate profile-generating script,\n"
        "and nothing in this notebook writes one."
    )
else:
    print("=== ETF Universe Profile ===")
    print(f"Written by {profile.source}")
    print(profile.summary())
```

## 6. Loader Options

The loader supports filtering by symbols and date range:

```python
# Specific symbols
spy_qqq = load_etfs(symbols=["SPY", "QQQ"])
print(f"SPY + QQQ only: {spy_qqq.shape}")
```

```python
# Date range
recent = load_etfs(start_date="2024-01-01")
print(f"2024 onwards: {recent.shape}")
```

```python
# Combined filters
filtered = load_etfs(
    symbols=["SPY", "QQQ", "IWM", "TLT", "GLD"], start_date="2020-01-01", end_date="2023-12-31"
)
print(f"5 ETFs, 2020-2023: {filtered.shape}")
```

## 7. Documentation

### Yahoo Finance
- [Yahoo Finance API (unofficial)](https://python-yahoofinance.readthedocs.io/)


### ETF Categories

| Category | Count | Description |
|----------|-------|-------------|
| `us_equity_broad` | 10 | Large, mid, small cap, equal weight |
| `us_equity_style` | 10 | Value, growth, momentum, dividend |
| `us_sectors` | 13 | SPDR sector ETFs + real estate |
| `international_developed` | 18 | EAFE, Europe, Japan, country ETFs |
| `emerging_markets` | 11 | EM broad + China, Brazil, India, etc. |
| `fixed_income` | 15 | Treasury, corporate, high yield, TIPS |
| `commodities` | 9 | Gold, silver, oil, broad commodity |
| `specialty` | 10 | Biotech, semiconductors, regional banks |
| `currency` | 4 | USD, EUR, JPY, GBP currency ETFs |

### Data Quality Notes
- Volume represents actual ETF trading volume
- Adjusted close accounts for dividends and splits
- Some ETFs have shorter history (check `first_date` in coverage)

## 8. Updating Data

To update with the latest data, re-run the download:

```python
# Update to latest available data
download_etf_data()

# Force full re-download
download_etf_data(force=True)
```

**Tip**: Update the `end` date in `config.yaml` before re-downloading
to extend the coverage period.

## Summary

| Item | Value |
|------|-------|
| Symbols | 100 ETFs across 9 categories |
| Frequency | Daily |
| Coverage | 2006-2025 |
| Provider | Yahoo Finance (free) |
| Config | `config.yaml` |
| Loader | `load_etfs(symbols, start_date, end_date)` |
| Profile | `$ML4T_DATA_PATH/etfs/market/etf_universe_profile.json` |

**Note**: This is the **candidate pool**. Chapter 6 filters to ~80 ETFs
based on liquidity, history, and correlation clustering.

Hiển thị toàn văn kèm ghi nguồn theo giấy phép của tài liệu gốc. Giấy phép: MIT

Bản tóm tắt này do tác nhân nghiên cứu của Stratmill biên soạn từ tài liệu gốc; đây không phải bản sao của tài liệu.