본문으로 건너뛰기
라이브러리 문서 전체

시계열 트레이딩 모델 학습·추론용 윈도우 데이터셋

코드 Machine Learning for Trading

요약

이 문서는 모델 학습과 롤링 추론을 위한 시계열 패널을 준비하는 데이터셋 유틸리티를 설명합니다. 학습 데이터셋은 설정 가능한 시퀀스 길이, 날짜 범위, 간격에 따라 특성, 선행 수익률, 변동성 척도, 마스크의 슬라이딩 시퀀스를 반환합니다. 날짜 범위는 사용 가능한 패널 날짜에 맞춰지며, 완전한 시퀀스를 제공할 수 없는 범위는 생성자가 거부합니다.

추론 데이터셋은 선택된 시간 인덱스에서 끝나는 윈도우를 생성하고, 각 예측 단계의 특성, 마스크, 해당 인덱스를 반환합니다. 정적 메타데이터 유틸리티는 정수형 자산 식별자를 할당하고, 선택적으로 자산 그룹과 거래 비용을 텐서로 인코딩할 수 있습니다. 이러한 인터페이스는 후속 모델에 시간 입력과 자산 맥락을 명시적으로 제공합니다. 발췌문에는 구현 동작과 입력 형태가 나와 있지만, 모델이나 트레이딩 결과, 검증 절차, 누수와 중첩 윈도우 의존성에 관한 논의는 없습니다. 이러한 속성은 호출자가 패널을 구성하고 사용하는 방식에 달려 있습니다.

핵심 아이디어

  • 학습 샘플은 특성, 수익률, 변동성 척도, 마스크 시퀀스를 담은 슬라이딩 윈도우입니다.
  • 날짜 범위와 간격이 포함할 학습 윈도우를 결정합니다.
  • 추론 윈도우는 지정된 시간 인덱스에서 끝나며 입력과 함께 해당 인덱스를 반환합니다.
  • 정적 메타데이터로 자산 식별 정보, 그룹, 거래 비용을 인코딩할 수 있습니다.
  • 데이터셋 유틸리티는 데이터 형태를 정의할 뿐, 트레이딩 모델의 성능이나 타당성을 입증하지 않습니다.

태그

전문
# dataset.py


```py
"""Torch datasets for DeePM-style windowed training and inference."""

from __future__ import annotations

from collections.abc import Sequence
from dataclasses import dataclass

import numpy as np
import pandas as pd
import torch
from torch.utils.data import Dataset

from .features import FeaturePanel


@dataclass(frozen=True, slots=True)
class StaticAssetMetadata:
    """Static per-asset metadata used as context."""

    assets: list[str]
    asset_ids: torch.Tensor  # (N,)
    group_ids: torch.Tensor | None  # (N,)
    costs: torch.Tensor | None  # (N, 1)


def build_static_metadata(
    assets: Sequence[str],
    *,
    asset_to_group: dict[str, str] | None = None,
    asset_to_cost_bps: dict[str, float] | None = None,
) -> StaticAssetMetadata:
    """Create tensors for asset id, group id, and costs."""
    assets_list = [str(a) for a in assets]
    n = len(assets_list)

    asset_ids = torch.arange(n, dtype=torch.long)

    group_ids: torch.Tensor | None = None
    if asset_to_group is not None:
        groups = [str(asset_to_group.get(a, "UNKNOWN")) for a in assets_list]
        unique_groups = {g: i for i, g in enumerate(sorted(set(groups)))}
        group_ids = torch.tensor([unique_groups[g] for g in groups], dtype=torch.long)

    costs: torch.Tensor | None = None
    if asset_to_cost_bps is not None:
        cost_vals = [float(asset_to_cost_bps.get(a, 0.0)) / 10000.0 for a in assets_list]
        costs = torch.tensor(cost_vals, dtype=torch.float32).unsqueeze(-1)

    return StaticAssetMetadata(
        assets=assets_list, asset_ids=asset_ids, group_ids=group_ids, costs=costs
    )


class DeepmWindowDataset(Dataset[tuple[torch.Tensor, torch.Tensor, torch.Tensor, torch.Tensor]]):
    """Sliding-window dataset for training.

    Each item returns (x_seq, y_seq, v_seq, m_seq) of shapes
    (L, N, F), (L, N), (L, N), (L, N).
    """

    def __init__(
        self,
        panel: FeaturePanel,
        *,
        seq_len: int,
        start_date: pd.Timestamp | None = None,
        end_date: pd.Timestamp | None = None,
        stride: int = 1,
    ) -> None:
        if seq_len <= 1:
            raise ValueError("seq_len must be > 1")

        self._panel = panel
        self.seq_len = int(seq_len)

        dates = panel.dates
        start_idx = 0
        end_idx_exclusive = len(dates)
        if start_date is not None:
            start_idx = int(dates.get_indexer([pd.Timestamp(start_date)], method="bfill")[0])
        if end_date is not None:
            end_idx_exclusive = (
                int(dates.get_indexer([pd.Timestamp(end_date)], method="ffill")[0]) + 1
            )

        t_max_start = (end_idx_exclusive - 1) - self.seq_len
        t_min_start = start_idx
        if t_max_start < t_min_start:
            raise ValueError("Date range too small for given seq_len")

        self.start_indices = np.arange(t_min_start, t_max_start + 1, stride, dtype=np.int64)

    def __len__(self) -> int:
        return int(self.start_indices.shape[0])

    def __getitem__(
        self, idx: int
    ) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor, torch.Tensor]:
        s = int(self.start_indices[idx])
        e = s + self.seq_len
        return (
            torch.from_numpy(self._panel.x[s:e]),
            torch.from_numpy(self._panel.y_fwd1[s:e]),
            torch.from_numpy(self._panel.vol_scale[s:e]),
            torch.from_numpy(self._panel.mask[s:e]),
        )


class RollingWindowInferenceDataset(Dataset[tuple[torch.Tensor, torch.Tensor, int]]):
    """Dataset for batched rolling-window inference."""

    def __init__(self, panel: FeaturePanel, *, seq_len: int, start_t: int) -> None:
        if seq_len <= 1:
            raise ValueError("seq_len must be > 1")
        if start_t < seq_len - 1:
            raise ValueError("start_t must be >= seq_len - 1")

        self._panel = panel
        self.seq_len = int(seq_len)
        self.times = np.arange(start_t, len(panel.dates) - 1, dtype=np.int64)

    def __len__(self) -> int:
        return int(self.times.shape[0])

    def __getitem__(self, idx: int) -> tuple[torch.Tensor, torch.Tensor, int]:
        t = int(self.times[idx])
        s = t - self.seq_len + 1
        e = t + 1
        return (
            torch.from_numpy(self._panel.x[s:e]),
            torch.from_numpy(self._panel.mask[s:e]),
            t,
        )

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.