Pular para o conteúdo
Todos os documentos da biblioteca

Como montar e avaliar um conjunto de dados multiativos de pesquisa ETF

Código Machine Learning for Trading

Resumo

Este documento descreve um universo diário de ativos candidatos ETF que abrange ações, renda fixa, commodities e moedas. Ele apresenta um fluxo de trabalho para baixar dados de mercado, carregá-los para análise, inspecionar a cobertura por símbolo e categoria e filtrar por símbolos ou intervalo de datas. Também explica que o grupo de candidatos se destina à seleção posterior de estratégias, na qual liquidez, histórico e agrupamento por correlação podem restringir o universo.

O perfil dos dados inclui volume real de negociação de ETF e preços de fechamento ajustados por dividendos e desdobramentos. A cobertura varia entre fundos, por isso recomenda-se que pesquisadores verifiquem a primeira data disponível para cada símbolo. O documento não apresenta estratégia de trading nem evidências de desempenho; seu valor está na descrição do conjunto de dados e no fluxo de pesquisa. Os dados vêm do Yahoo Finance, e a cobertura e a composição do universo informadas descrevem este conjunto de dados específico, sem garantir históricos completos ou uniformes. Como em outros conjuntos de dados retrospectivos de mercado, pesquisadores devem verificar a cobertura e a qualidade dos dados antes de usá-los em um estudo de estratégia.

Ideias principais

  • O universo de candidatos abrange várias classes de ativos e categorias ETF em frequência diária.
  • Pesquisadores podem carregar subconjuntos por símbolo e intervalo de datas para analisar estratégias.
  • Liquidez, histórico e agrupamento por correlação devem orientar a seleção posterior do universo.
  • O volume de ETF é informado como volume negociado, e os preços de fechamento ajustados consideram dividendos e desdobramentos.
  • O histórico disponível varia entre fundos, portanto a cobertura deve ser verificada por símbolo.

Tags

Texto completo
# dataset_card.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: -all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # ETF Universe Dataset
#
# 100 diversified ETFs across 9 categories for momentum and cross-asset strategies.
#
# | Property | Value |
# |----------|-------|
# | **Provider** | Yahoo Finance |
# | **Asset Class** | Multi-asset (Equity, Fixed Income, Commodities, Currency) |
# | **Frequency** | Daily |
# | **Symbols** | 100 ETFs |
# | **Coverage** | 2006-2025 |
# | **Size** | ~16 MB |
# | **API Key** | None (free) |
# | **Loader** | `load_etfs()` |

# %%
"""ETF Universe - download, explore, and update workflow."""

import json
from pathlib import Path

import polars as pl
import yaml

# %% [markdown]
# ## 1. Configuration
#
# The ETF universe is defined in `config.yaml`. This is the **candidate pool** -
# strategy definition (Chapter 6) filters this down based on liquidity, history,
# and correlation clustering.

# %%
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]

print("=== ETF Configuration ===")
print(f"Provider: {etf_config['provider']}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Frequency: {etf_config['frequency']}")
print(f"\nCategories ({len(etf_config['tickers'])}):")
for category, info in etf_config["tickers"].items():
    symbols = info["symbols"]
    print(f"  {category}: {len(symbols)} ETFs")

total_etfs = sum(len(info["symbols"]) for info in etf_config["tickers"].values())
print(f"\nTotal: {total_etfs} ETFs")

# %% [markdown]
# ## 2. API Key Setup
#
# **No API key required.** Yahoo Finance data is free and publicly accessible.
#
# The `ml4t-data` library handles rate limiting automatically to avoid
# being blocked by Yahoo Finance.

# %%
print("Yahoo Finance requires no API key - data is publicly available.")

# %% [markdown]
# ## 3. Download Data
#
# The download uses the `ml4t-data` library which handles:
# - Rate limiting (1 second delay between batches)
# - Retry logic for failed requests
# - Consistent schema output
#
# **Note**: First-time download takes ~2-3 minutes for 100 ETFs.


# %%
def download_etf_data(dry_run: bool = False, force: bool = False, symbols: list[str] | None = None):
    """Download ETF data from Yahoo Finance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from ml4t.data.providers import YahooFinanceProvider

    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    etf_config = config["etfs"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in etf_config["tickers"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "etfs" / "market"
    output_path = output_dir / "etf_universe.parquet"

    print("=== ETF Download ===")
    print(f"Symbols: {len(symbols)}")
    print(f"Date range: {etf_config['start']} to {etf_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for i, symbol in enumerate(symbols, 1):
            print(f"  {i:3}. {symbol}")
        return

    # Check existing data
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        existing_symbols = set(existing["symbol"].unique().to_list())
        missing = [s for s in symbols if s not in existing_symbols]
        if not missing:
            print(f"\nAll {len(symbols)} ETFs already downloaded.")
            print("Use force=True to re-download.")
            return existing
        print(f"Found {len(existing_symbols)} existing, downloading {len(missing)} missing...")
        symbols = missing

    # Initialize provider and download
    provider = YahooFinanceProvider()
    print(f"\nDownloading {len(symbols)} ETFs...")

    etf_data = provider.fetch_batch_ohlcv(
        symbols=symbols,
        start=etf_config["start"],
        end=etf_config["end"],
        frequency="daily",
        chunk_size=50,
        delay_seconds=1.0,
    )

    # Combine with existing data if applicable
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        etf_data = pl.concat([existing, etf_data])

    # Save
    output_dir.mkdir(parents=True, exist_ok=True)
    etf_data.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(etf_data):,}")
    print(f"Symbols: {etf_data['symbol'].n_unique()}")
    print(f"Date range: {etf_data['timestamp'].min()} to {etf_data['timestamp'].max()}")
    print(f"Saved to: {output_path}")

    return etf_data


# %% [markdown]
# ### Download All ETFs

# %%
# Uncomment to download all ETF data
# download_etf_data()

# %% [markdown]
# ### Dry Run (Preview)
#
# See what would be downloaded without actually downloading:

# %%
download_etf_data(dry_run=True)

# %% [markdown]
# ## 4. Load and Explore
#
# Once downloaded, use the loader throughout the book:

# %%
from data import load_etfs

# Load all ETF data
df = load_etfs()

print(f"Shape: {df.shape}")
print(f"Symbols: {df['symbol'].n_unique()}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")

# %%
# Schema
df.schema

# %%
# Preview
df.head(10)

# %% [markdown]
# ### Coverage by Symbol

# %%
# Coverage and basic stats by symbol
coverage = (
    df.group_by("symbol")
    .agg(
        pl.col("timestamp").min().alias("first_date"),
        pl.col("timestamp").max().alias("last_date"),
        pl.len().alias("n_bars"),
        pl.col("volume").mean().alias("avg_daily_volume"),
    )
    .sort("avg_daily_volume", descending=True)
)
coverage.head(20)

# %% [markdown]
# ### Category Summary

# %%
# Build category mapping from config
category_map = {}
for category, info in etf_config["tickers"].items():
    for symbol in info["symbols"]:
        category_map[symbol] = category

df_with_cat = df.with_columns(pl.col("symbol").replace(category_map).alias("category"))

category_summary = (
    df_with_cat.group_by("category")
    .agg(
        pl.col("symbol").n_unique().alias("n_symbols"),
        pl.col("timestamp").min().alias("earliest"),
        pl.col("timestamp").max().alias("latest"),
        pl.col("volume").mean().alias("avg_volume"),
    )
    .sort("n_symbols", descending=True)
)
category_summary

# %% [markdown]
# ## 5. Data Profile
#
# Profiles document the dataset structure, statistics, and quality metrics.
# They are stored alongside the data files.

# %%
from ml4t.data.storage.data_profile import load_profile

from utils import ML4T_DATA_PATH

profile_path = ML4T_DATA_PATH / "etfs" / "market" / "etf_universe_profile.json"
profile = load_profile(profile_path)

if profile is None:
    print(f"No profile at {profile_path}")
    print(
        "Profiles are written next to the data by whatever builds the dataset - the\n"
        "download script in this directory, or the ml4t-data loader it drives - through\n"
        "ml4t.data.storage.data_profile. There is no separate profile-generating script,\n"
        "and nothing in this notebook writes one."
    )
else:
    print("=== ETF Universe Profile ===")
    print(f"Written by {profile.source}")
    print(profile.summary())

# %% [markdown]
# ## 6. Loader Options
#
# The loader supports filtering by symbols and date range:

# %%
# Specific symbols
spy_qqq = load_etfs(symbols=["SPY", "QQQ"])
print(f"SPY + QQQ only: {spy_qqq.shape}")

# %%
# Date range
recent = load_etfs(start_date="2024-01-01")
print(f"2024 onwards: {recent.shape}")

# %%
# Combined filters
filtered = load_etfs(
    symbols=["SPY", "QQQ", "IWM", "TLT", "GLD"], start_date="2020-01-01", end_date="2023-12-31"
)
print(f"5 ETFs, 2020-2023: {filtered.shape}")

# %% [markdown]
# ## 7. Documentation
#
# ### Yahoo Finance
# - [Yahoo Finance API (unofficial)](https://python-yahoofinance.readthedocs.io/)
# - Rate limits: ~2000 requests/hour (handled by ml4t-data)
#
# ### ETF Categories
#
# | Category | Count | Description |
# |----------|-------|-------------|
# | `us_equity_broad` | 10 | Large, mid, small cap, equal weight |
# | `us_equity_style` | 10 | Value, growth, momentum, dividend |
# | `us_sectors` | 13 | SPDR sector ETFs + real estate |
# | `international_developed` | 18 | EAFE, Europe, Japan, country ETFs |
# | `emerging_markets` | 11 | EM broad + China, Brazil, India, etc. |
# | `fixed_income` | 15 | Treasury, corporate, high yield, TIPS |
# | `commodities` | 9 | Gold, silver, oil, broad commodity |
# | `specialty` | 10 | Biotech, semiconductors, regional banks |
# | `currency` | 4 | USD, EUR, JPY, GBP currency ETFs |
#
# ### Data Quality Notes
# - Volume represents actual ETF trading volume
# - Adjusted close accounts for dividends and splits
# - Some ETFs have shorter history (check `first_date` in coverage)

# %% [markdown]
# ## 8. Updating Data
#
# To update with the latest data, re-run the download:
#
# ```python
# # Update to latest available data
# download_etf_data()
#
# # Force full re-download
# download_etf_data(force=True)
# ```
#
# **Tip**: Update the `end` date in `config.yaml` before re-downloading
# to extend the coverage period.

# %% [markdown]
# ## Summary
#
# | Item | Value |
# |------|-------|
# | Symbols | 100 ETFs across 9 categories |
# | Frequency | Daily |
# | Coverage | 2006-2025 |
# | Provider | Yahoo Finance (free) |
# | Config | `config.yaml` |
# | Loader | `load_etfs(symbols, start_date, end_date)` |
# | Profile | `$ML4T_DATA_PATH/etfs/market/etf_universe_profile.json` |
#
# **Note**: This is the **candidate pool**. Chapter 6 filters to ~80 ETFs
# based on liquidity, history, and correlation clustering.

```

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.