본문으로 건너뛰기
라이브러리 문서 전체

펀딩 및 베이시스 연구용 암호화폐 무기한 선물 프리미엄 데이터

노트북 Machine Learning for Trading

요약

이 문서는 암호화폐 무기한 선물과 프리미엄 지수를 함께 연구하기 위한 데이터셋과 워크플로를 설명합니다. 설정된 종목군의 시간별 OHLCV 관측값과 8시간 간격의 프리미엄 값을 다루며, 다운로드, 로드, 필터링, 기본 프로파일링 절차를 안내합니다. 프리미엄 지수는 무기한 선물 가격과 현물 가격 간 상대적 차이로 정의되며, 부호로 무기한 계약이 현물보다 높거나 낮게 거래되는지 나타냅니다. 이 문서는 데이터를 펀딩비 차익거래와 프리미엄 평균회귀 연구의 입력으로 제시하며, 검증된 전략을 입증하지는 않습니다.

데이터는 Binance 공개 피드에서 가져오며 거래소별 거래량을 포함합니다. 토큰마다 커버리지가 다릅니다. 기존 자산은 더 긴 과거 데이터가 있고 신규 상장 자산은 더 늦게 시작하며, 심볼 하나는 이름이 바뀌었습니다. 프리미엄 관측값은 펀딩 정산 간격에 맞춰집니다. 연구를 설계할 때 이런 커버리지 차이, 특정 거래소로의 집중, 그리고 베이시스 지표와 실현 펀딩 또는 거래 가능한 수익률의 차이를 고려해야 합니다.

핵심 아이디어

  • 데이터셋은 시간별 무기한 선물 OHLCV과 8시간 간격 프리미엄 지수 관측값을 짝지어 제공합니다.
  • 프리미엄은 무기한 계약과 현물 간 상대 가격 차이를 측정합니다.
  • 프리미엄이 양수인지 음수인지에 따라 무기한 선물이 현물보다 높은지 낮은지 알 수 있습니다.
  • 이 데이터는 펀딩비 차익거래와 프리미엄 평균회귀 연구에 쓸 수 있지만 수익성을 입증하지는 않습니다.
  • 자산별 데이터 시작 시점이 다르며 거래량은 Binance 데이터만 반영합니다.

태그

전문
# Crypto Premium Index Dataset


# Crypto Premium Index Dataset

Perpetual futures OHLCV and premium index data for funding rate arbitrage strategy.

| Property | Value |
|----------|-------|
| **Provider** | Binance Public API |
| **Asset Class** | Cryptocurrency |
| **Frequency** | 1h (OHLCV), 8h (premium) |
| **Symbols** | 20 perpetual futures |
| **Coverage** | 2020-2025 |
| **Size** | ~70 MB |
| **API Key** | None (free) |
| **Loader** | `load_crypto_perps()`, `load_crypto_premium()` |

```python
"""Crypto Premium Index - download, explore, and update workflow."""

from pathlib import Path

import polars as pl
import yaml
```

## 1. Configuration

The crypto universe is defined in `config.yaml`. Organized by market segment:
major cryptocurrencies, DeFi tokens, and Layer 1 blockchains.

```python
# Load and display configuration
config_path = Path("config.yaml")
config = yaml.safe_load(config_path.read_text())
crypto_config = config["crypto"]

print("=== Crypto Configuration ===")
print(f"Provider: {crypto_config['provider']}")
print(f"Market: {crypto_config['market']}")
print(f"Date range: {crypto_config['start']} to {crypto_config['end']}")
print(f"Premium interval: {crypto_config['interval']}")
print("\nCategories:")
for category, info in crypto_config["symbols"].items():
    symbols = info["symbols"]
    print(f"  {category}: {len(symbols)} tokens - {info['description']}")
    print(f"    {', '.join(symbols[:5])}{'...' if len(symbols) > 5 else ''}")

total_symbols = sum(len(info["symbols"]) for info in crypto_config["symbols"].values())
print(f"\nTotal: {total_symbols} symbols")
```

## 2. API Key Setup

**No API key required.** Binance Public API provides free access to historical data
through data.binance.vision.

The `ml4t-data` library handles rate limiting automatically.

```python
print("Binance Public API requires no API key - data is freely available.")
print("Source: data.binance.vision")
```

## 3. Download Data

The download uses the `ml4t-data` library which handles:
- Rate limiting
- Data validation
- Consistent schema output

Two types of data are available:
- **OHLCV** (hourly): Price and volume for perpetual futures
- **Premium Index** (8-hourly): Basis between perpetual and spot prices

```python
def download_crypto_ohlcv(
    dry_run: bool = False, force: bool = False, symbols: list[str] | None = None
):
    """Download crypto perpetual futures OHLCV from Binance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    crypto_config = config["crypto"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in crypto_config["symbols"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "crypto" / "market"
    output_path = output_dir / "perps_1h.parquet"

    print("=== Crypto OHLCV Download ===")
    print(f"Symbols: {len(symbols)}")
    print("Frequency: 1h")
    print(f"Date range: {crypto_config['start']} to {crypto_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for symbol in symbols:
            print(f"  {symbol}")
        return

    # Check existing
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        print(f"\nData already exists ({len(existing):,} rows).")
        print("Use force=True to re-download.")
        return existing

    # Initialize provider
    from ml4t.data.providers import BinancePublicProvider

    provider = BinancePublicProvider(market="spot")

    # Download each symbol
    all_data = []
    print(f"\nDownloading {len(symbols)} symbols...")
    for symbol in symbols:
        print(f"  {symbol}...", end=" ", flush=True)
        try:
            df = provider.fetch_ohlcv(
                symbol=symbol,
                start=crypto_config["start"],
                end=crypto_config["end"],
                frequency="hourly",
            )
            df = df.with_columns(pl.lit(symbol).alias("symbol"))
            all_data.append(df)
            print(f"OK ({len(df):,} rows)")
        except Exception as e:
            print(f"ERROR: {e}")

    if not all_data:
        raise RuntimeError("No data downloaded!")

    # Combine and save
    output_dir.mkdir(parents=True, exist_ok=True)
    combined = pl.concat(all_data)
    combined.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(combined):,}")
    print(f"Symbols: {combined['symbol'].n_unique()}")
    print(f"Saved to: {output_path}")

    return combined


def download_crypto_premium(
    dry_run: bool = False, force: bool = False, symbols: list[str] | None = None
):
    """Download crypto premium index from Binance.

    Args:
        dry_run: If True, show what would be downloaded without doing it
        force: If True, re-download even if data exists
        symbols: Specific symbols to download (default: all from config)
    """
    from utils import ML4T_DATA_PATH

    # Load config
    config = yaml.safe_load(config_path.read_text())
    crypto_config = config["crypto"]

    # Flatten symbols list
    if symbols is None:
        symbols = []
        for category_info in crypto_config["symbols"].values():
            symbols.extend(category_info["symbols"])

    output_dir = ML4T_DATA_PATH / "crypto" / "market"
    output_path = output_dir / "premium_index_8h.parquet"

    print("=== Crypto Premium Index Download ===")
    print(f"Symbols: {len(symbols)}")
    print("Frequency: 8h (funding rate interval)")
    print(f"Date range: {crypto_config['start']} to {crypto_config['end']}")
    print(f"Output: {output_path}")

    if dry_run:
        print("\n[DRY RUN] Would download:")
        for symbol in symbols:
            print(f"  {symbol}")
        return

    # Check existing
    if output_path.exists() and not force:
        existing = pl.read_parquet(output_path)
        print(f"\nData already exists ({len(existing):,} rows).")
        print("Use force=True to re-download.")
        return existing

    # Initialize provider (futures market for premium index)
    from ml4t.data.providers import BinancePublicProvider

    provider = BinancePublicProvider(market="futures")

    print(f"\nDownloading premium index for {len(symbols)} symbols...")
    premium_data = provider.fetch_premium_index_multi(
        symbols=symbols,
        start=crypto_config["start"],
        end=crypto_config["end"],
        interval="8h",
    )

    # Save
    output_dir.mkdir(parents=True, exist_ok=True)
    premium_data.write_parquet(output_path)

    print("\n=== Complete ===")
    print(f"Total rows: {len(premium_data):,}")
    print(f"Symbols: {premium_data['symbol'].n_unique()}")
    print(f"Saved to: {output_path}")

    return premium_data
```

### Download OHLCV Data

```python
# Uncomment to download OHLCV data
# download_crypto_ohlcv()
```

### Download Premium Index

```python
# Uncomment to download premium index
# download_crypto_premium()
```

### Dry Run (Preview)

```python
download_crypto_ohlcv(dry_run=True)
```

## 4. Load and Explore

Once downloaded, use the loaders throughout the book:

```python
from data import load_crypto_perps, load_crypto_premium
```

### Perpetual Futures OHLCV

```python
# Load hourly OHLCV data
perps = load_crypto_perps()

print(f"Shape: {perps.shape}")
print(f"Symbols: {perps['symbol'].n_unique()}")
print(f"Date range: {perps['timestamp'].min()} to {perps['timestamp'].max()}")
print(f"Memory: {perps.estimated_size('mb'):.1f} MB")
```

```python
# Schema
perps.schema
```

```python
# Preview
perps.head(10)
```

```python
# Volume by symbol (USD notional)
volume_by_symbol = (
    perps.group_by("symbol")
    .agg(
        (pl.col("volume") * pl.col("close")).sum().alias("total_volume_usd"),
        pl.len().alias("n_observations"),
        pl.col("timestamp").min().alias("first_date"),
        pl.col("timestamp").max().alias("last_date"),
    )
    .sort("total_volume_usd", descending=True)
)
volume_by_symbol
```

### Premium Index

```python
# Load 8-hourly premium index data
premium = load_crypto_premium()

print(f"Shape: {premium.shape}")
print(f"Symbols: {premium['symbol'].n_unique()}")
print(f"Date range: {premium['timestamp'].min()} to {premium['timestamp'].max()}")
print(f"Memory: {premium.estimated_size('mb'):.1f} MB")
```

```python
# Preview
premium.head(10)
```

```python
# Premium statistics by symbol
# Premium index captures basis between perpetual and spot
premium_stats = (
    premium.group_by("symbol")
    .agg(
        pl.col("premium_index_close").mean().alias("mean_premium"),
        pl.col("premium_index_close").std().alias("std_premium"),
        pl.col("premium_index_close").min().alias("min_premium"),
        pl.col("premium_index_close").max().alias("max_premium"),
    )
    .sort("mean_premium", descending=True)
)
premium_stats
```

## 5. Data Profile

Profiles document the dataset structure, statistics, and quality metrics.

```python
from ml4t.data.storage.data_profile import load_profile

from utils import ML4T_DATA_PATH
from utils.paths import display_path

for dataset, filename in [
    ("OHLCV", "perps_1h_profile.json"),
    ("Premium", "premium_index_8h_profile.json"),
]:
    profile_path = ML4T_DATA_PATH / "crypto" / "market" / filename
    profile = load_profile(profile_path)
    if profile is None:
        print(f"No profile at {display_path(profile_path)}")
        print(
            "Profiles are written next to the data by whatever builds the dataset, through\n"
            "ml4t.data.storage.data_profile. There is no separate profile-generating\n"
            "script, and nothing in this notebook writes one.\n"
        )
    else:
        print(f"=== Crypto {dataset} Profile ===")
        print(f"Written by {profile.source}")
        print(profile.summary())
        print()
```

## 6. Loader Options

The loaders support filtering by symbols and date range:

```python
# Specific symbols
btc_eth = load_crypto_perps(symbols=["BTCUSDT", "ETHUSDT"])
print(f"BTC + ETH only: {btc_eth.shape}")
```

```python
# Date range
recent = load_crypto_premium(start_date="2024-01-01")
print(f"Premium 2024+: {recent.shape}")
```

```python
# Combined filters
filtered = load_crypto_perps(
    symbols=["BTCUSDT", "ETHUSDT", "SOLUSDT"], start_date="2023-01-01", end_date="2023-12-31"
)
print(f"3 symbols, 2023: {filtered.shape}")
```

## 7. Documentation

### Binance Public API
- [Binance Public Data](https://data.binance.vision/)
- [API Documentation](https://binance-docs.github.io/apidocs/spot/en/)

### Premium Index

The premium index measures the basis between perpetual futures and spot prices:

$$\text{Premium} = \frac{P_{perp} - P_{spot}}{P_{spot}}$$

Key properties:
- **Positive premium**: Perpetual trades at premium (bullish sentiment)
- **Negative premium**: Perpetual trades at discount (bearish sentiment)
- **Funding rate**: Derived from premium, settles every 8 hours

### Data Quality Notes
- Volume represents Binance exchange volume only
- BTC/ETH/major alts: Data from Jan 2020 (6 years history)
- Newer tokens (APT, SUI, INJ, ARB, OP): Data from listing date (2022-2023)
- MATICUSDT renamed to POLUSDT in Sept 2024 (data ends there)
- 8-hour intervals align with funding rate settlement times (00:00, 08:00, 16:00 UTC)

## 8. Updating Data

To update with the latest data, re-run the download:

```python
# Update OHLCV data
download_crypto_ohlcv()

# Update premium index
download_crypto_premium()

# Force full re-download
download_crypto_ohlcv(force=True)
download_crypto_premium(force=True)
```

**Tip**: Update the `end` date in `config.yaml` before re-downloading.

## Summary

| Item | Value |
|------|-------|
| Symbols | 20 perpetual futures (major, DeFi, L1) |
| Frequencies | 1h OHLCV, 8h premium index |
| Coverage | 2020-2025 (6 years for BTC/ETH) |
| Provider | Binance Public (free) |
| Config | `config.yaml` |
| Loaders | `load_crypto_perps()`, `load_crypto_premium()` |

**Use case**: Funding rate arbitrage strategy exploiting premium mean reversion.

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.