Saltar al contenido
Todos los documentos de la biblioteca

Creación y evaluación de un conjunto de datos ETF multiactivo

Código Machine Learning for Trading

Resumen

Este documento describe un universo diario de candidatos ETF que abarca acciones, renta fija, materias primas y divisas. Expone un flujo de trabajo para descargar datos de mercado, cargarlos para analizarlos, inspeccionar la cobertura por símbolo y categoría, y filtrar por símbolos o intervalo de fechas. También explica que el conjunto de candidatos está pensado para una selección posterior de estrategias, en la que la liquidez, el historial y la agrupación por correlación pueden reducir el universo.

El perfil de datos incluye el volumen real de negociación de ETF y precios de cierre ajustados que tienen en cuenta dividendos y desdoblamientos. La cobertura varía según el fondo, por lo que se recomienda comprobar la primera fecha disponible de cada símbolo. El documento no aporta ninguna estrategia de trading ni evidencia de rendimiento; su valor reside en describir el conjunto de datos y el flujo de investigación. Los datos proceden de Yahoo Finance y la cobertura y composición del universo indicadas describen este conjunto de datos específico, no garantizan historiales completos o uniformes. Como ocurre con otros conjuntos de datos de mercado retrospectivos, conviene verificar la cobertura y la calidad de los datos antes de usarlos en un estudio de estrategias.

Ideas clave

  • El universo de candidatos abarca varias clases de activos y categorías ETF con frecuencia diaria.
  • Puedes cargar subconjuntos por símbolo e intervalo de fechas para analizar estrategias.
  • La liquidez, el historial y la agrupación por correlación están pensados para orientar la selección posterior del universo.
  • El volumen de ETF se presenta como volumen negociado, y los precios de cierre ajustados tienen en cuenta dividendos y desdoblamientos.
  • El historial disponible varía entre fondos, así que conviene comprobar la cobertura de cada símbolo.

Etiquetas

Texto completo
# dataset_card.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: -all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # ETF Universe Dataset
#
# 100 diversified ETFs across 9 categories for momentum and cross-asset strategies.
#
# | Property | Value |
# |----------|-------|
# | **Provider** | Yahoo Finance |
# | **Asset Class** | Multi-asset (Equity, Fixed Income, Commodities, Currency) |
# | **Frequency** | Daily |
# | **Symbols** | 100 ETFs |
# | **Coverage** | 2006-2025 |
# | **Size** | ~16 MB |
# | **API Key** | None (free) |
# | **Loader** | `load_etfs()` |

# %%
"""ETF Universe - download, explore, and update workflow."""

import json
from pathlib import Path

import polars as pl
import yaml

# %% [markdown]
# ## 1. Configuration
#
# The ETF universe is defined in `config.yaml`. This is the **candidate pool** -
# strategy definition (Chapter 6) filters this down based on liquidity, history,
# and correlation clustering.

# %%
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
etf_config = config["etfs"]

print("=== ETF Configuration ===")
print(f"Provider: {etf_config['provider']}")
print(f"Date range: {etf_config['start']} to {etf_config['end']}")
print(f"Frequency: {etf_config['frequency']}")
print(f"\nCategories ({len(etf_config['tickers'])}):")
for category, info in etf_config["tickers"].items():
    symbols = info["symbols"]
    print(f"  {category}: {len(symbols)} ETFs")

total_etfs = sum(len(info["symbols"]) for info in etf_config["tickers"].values())
print(f"\nTotal: {total_etfs} ETFs")

# %% [markdown]
# ## 2. API Key Setup
#
# **No API key required.** Yahoo Finance data is free and publicly accessible.
#
# The `ml4t-data` library handles rate limiting automatically to avoid
# being blocked by Yahoo Finance.

# %%
print("Yahoo Finance requires no API key - data is publicly available.")

# %% [markdown]
# ## 3. Download Data
#
# The download uses the `ml4t-data` library which handles:
# - Rate limiting (1 second delay between batches)
# - Retry logic for failed requests
# - Consistent schema output
#
# **Note**: First-time download takes ~2-3 minutes for 100 ETFs.


# %%
def download_etf_data(dry_run: bool = False, force: bool = False, symbols: list[str] | None = None):
    """Download ETF data from Yahoo Finance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from ml4t.data.providers import YahooFinanceProvider

    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    etf_config = config["etfs"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in etf_config["tickers"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "etfs" / "market"
    output_path = output_dir / "etf_universe.parquet"

    print("=== ETF Download ===")
    print(f"Symbols: {len(symbols)}")
    print(f"Date range: {etf_config['start']} to {etf_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for i, symbol in enumerate(symbols, 1):
            print(f"  {i:3}. {symbol}")
        return

    # Check existing data
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        existing_symbols = set(existing["symbol"].unique().to_list())
        missing = [s for s in symbols if s not in existing_symbols]
        if not missing:
            print(f"\nAll {len(symbols)} ETFs already downloaded.")
            print("Use force=True to re-download.")
            return existing
        print(f"Found {len(existing_symbols)} existing, downloading {len(missing)} missing...")
        symbols = missing

    # Initialize provider and download
    provider = YahooFinanceProvider()
    print(f"\nDownloading {len(symbols)} ETFs...")

    etf_data = provider.fetch_batch_ohlcv(
        symbols=symbols,
        start=etf_config["start"],
        end=etf_config["end"],
        frequency="daily",
        chunk_size=50,
        delay_seconds=1.0,
    )

    # Combine with existing data if applicable
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        etf_data = pl.concat([existing, etf_data])

    # Save
    output_dir.mkdir(parents=True, exist_ok=True)
    etf_data.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(etf_data):,}")
    print(f"Symbols: {etf_data['symbol'].n_unique()}")
    print(f"Date range: {etf_data['timestamp'].min()} to {etf_data['timestamp'].max()}")
    print(f"Saved to: {output_path}")

    return etf_data


# %% [markdown]
# ### Download All ETFs

# %%
# Uncomment to download all ETF data
# download_etf_data()

# %% [markdown]
# ### Dry Run (Preview)
#
# See what would be downloaded without actually downloading:

# %%
download_etf_data(dry_run=True)

# %% [markdown]
# ## 4. Load and Explore
#
# Once downloaded, use the loader throughout the book:

# %%
from data import load_etfs

# Load all ETF data
df = load_etfs()

print(f"Shape: {df.shape}")
print(f"Symbols: {df['symbol'].n_unique()}")
print(f"Date range: {df['timestamp'].min()} to {df['timestamp'].max()}")
print(f"Memory: {df.estimated_size('mb'):.1f} MB")

# %%
# Schema
df.schema

# %%
# Preview
df.head(10)

# %% [markdown]
# ### Coverage by Symbol

# %%
# Coverage and basic stats by symbol
coverage = (
    df.group_by("symbol")
    .agg(
        pl.col("timestamp").min().alias("first_date"),
        pl.col("timestamp").max().alias("last_date"),
        pl.len().alias("n_bars"),
        pl.col("volume").mean().alias("avg_daily_volume"),
    )
    .sort("avg_daily_volume", descending=True)
)
coverage.head(20)

# %% [markdown]
# ### Category Summary

# %%
# Build category mapping from config
category_map = {}
for category, info in etf_config["tickers"].items():
    for symbol in info["symbols"]:
        category_map[symbol] = category

df_with_cat = df.with_columns(pl.col("symbol").replace(category_map).alias("category"))

category_summary = (
    df_with_cat.group_by("category")
    .agg(
        pl.col("symbol").n_unique().alias("n_symbols"),
        pl.col("timestamp").min().alias("earliest"),
        pl.col("timestamp").max().alias("latest"),
        pl.col("volume").mean().alias("avg_volume"),
    )
    .sort("n_symbols", descending=True)
)
category_summary

# %% [markdown]
# ## 5. Data Profile
#
# Profiles document the dataset structure, statistics, and quality metrics.
# They are stored alongside the data files.

# %%
from ml4t.data.storage.data_profile import load_profile

from utils import ML4T_DATA_PATH

profile_path = ML4T_DATA_PATH / "etfs" / "market" / "etf_universe_profile.json"
profile = load_profile(profile_path)

if profile is None:
    print(f"No profile at {profile_path}")
    print(
        "Profiles are written next to the data by whatever builds the dataset - the\n"
        "download script in this directory, or the ml4t-data loader it drives - through\n"
        "ml4t.data.storage.data_profile. There is no separate profile-generating script,\n"
        "and nothing in this notebook writes one."
    )
else:
    print("=== ETF Universe Profile ===")
    print(f"Written by {profile.source}")
    print(profile.summary())

# %% [markdown]
# ## 6. Loader Options
#
# The loader supports filtering by symbols and date range:

# %%
# Specific symbols
spy_qqq = load_etfs(symbols=["SPY", "QQQ"])
print(f"SPY + QQQ only: {spy_qqq.shape}")

# %%
# Date range
recent = load_etfs(start_date="2024-01-01")
print(f"2024 onwards: {recent.shape}")

# %%
# Combined filters
filtered = load_etfs(
    symbols=["SPY", "QQQ", "IWM", "TLT", "GLD"], start_date="2020-01-01", end_date="2023-12-31"
)
print(f"5 ETFs, 2020-2023: {filtered.shape}")

# %% [markdown]
# ## 7. Documentation
#
# ### Yahoo Finance
# - [Yahoo Finance API (unofficial)](https://python-yahoofinance.readthedocs.io/)
# - Rate limits: ~2000 requests/hour (handled by ml4t-data)
#
# ### ETF Categories
#
# | Category | Count | Description |
# |----------|-------|-------------|
# | `us_equity_broad` | 10 | Large, mid, small cap, equal weight |
# | `us_equity_style` | 10 | Value, growth, momentum, dividend |
# | `us_sectors` | 13 | SPDR sector ETFs + real estate |
# | `international_developed` | 18 | EAFE, Europe, Japan, country ETFs |
# | `emerging_markets` | 11 | EM broad + China, Brazil, India, etc. |
# | `fixed_income` | 15 | Treasury, corporate, high yield, TIPS |
# | `commodities` | 9 | Gold, silver, oil, broad commodity |
# | `specialty` | 10 | Biotech, semiconductors, regional banks |
# | `currency` | 4 | USD, EUR, JPY, GBP currency ETFs |
#
# ### Data Quality Notes
# - Volume represents actual ETF trading volume
# - Adjusted close accounts for dividends and splits
# - Some ETFs have shorter history (check `first_date` in coverage)

# %% [markdown]
# ## 8. Updating Data
#
# To update with the latest data, re-run the download:
#
# ```python
# # Update to latest available data
# download_etf_data()
#
# # Force full re-download
# download_etf_data(force=True)
# ```
#
# **Tip**: Update the `end` date in `config.yaml` before re-downloading
# to extend the coverage period.

# %% [markdown]
# ## Summary
#
# | Item | Value |
# |------|-------|
# | Symbols | 100 ETFs across 9 categories |
# | Frequency | Daily |
# | Coverage | 2006-2025 |
# | Provider | Yahoo Finance (free) |
# | Config | `config.yaml` |
# | Loader | `load_etfs(symbols, start_date, end_date)` |
# | Profile | `$ML4T_DATA_PATH/etfs/market/etf_universe_profile.json` |
#
# **Note**: This is the **candidate pool**. Chapter 6 filters to ~80 ETFs
# based on liquidity, history, and correlation clustering.

```

Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT

Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.