본문으로 건너뛰기
라이브러리 문서 전체

자산 간 포트폴리오 리스크 비중을 산출하는 DeePM 신경망

코드 Machine Learning for Trading

요약

이 문서는 자산별 시계열과 맥락 정보를 상한과 하한이 있는 포트폴리오 리스크 비중으로 변환하는 신경망 정책 네트워크의 구조를 설명합니다. 자산별 기본 구조는 맥락 임베딩, 특성별 조절, 변수 선택, LSTM, 인과적 시간 어텐션을 결합합니다. 이후 지연된 횡단면 어텐션으로 자산 간 관계를 처리하며, 선택적으로 거시경제 인접 그래프로 범위를 제한한 어텐션을 적용할 수 있습니다.

정적 맥락에는 자산 정체성, 그룹 소속, 거래 비용을 담을 수 있고, 마스크는 이용 불가능한 자산을 처리합니다. 출력 헤드는 쌍곡탄젠트를 사용하고 자산 마스크를 적용해 유효한 자산에 대해 -1과 1 사이의 비중을 산출합니다. 이 문서는 구조 구성 요소와 유효한 키가 없는 어텐션 행을 방지하는 장치를 설명하지만, 학습 절차, 트레이딩 결과, 벤치마크 비교 또는 이 구조가 포트폴리오 성과를 개선한다는 근거는 제시하지 않습니다. 따라서 이 설명만으로는 모델 설계를 이해할 수 있을 뿐 실증적 효과를 판단할 수 없습니다.

핵심 아이디어

  • 자산 맥락은 입력 특성과 순환 은닉 상태 모두에 영향을 줄 수 있습니다.
  • 변수 선택은 맥락에 따라 달라지는 가중치로 여러 특성을 결합합니다.
  • 순환 계층과 인과적 어텐션 계층으로 시간 패턴을 표현합니다.
  • 지연된 횡단면 어텐션과 선택적 그래프 마스크로 자산 간 관계를 모델링합니다.
  • 제한된 출력 헤드는 자산 가용성을 반영하면서 부호가 있는 리스크 비중을 산출합니다.

태그

전문
# model.py


```py
"""PyTorch implementation of the DeePM deep portfolio manager.

Architecture components:
1. Per-asset temporal backbone (shared weights): FiLM conditioning, variable
   selection, LSTM, temporal self-attention.
2. Cross-sectional attention with Directed Delay for causality.
3. Macroeconomic graph prior as adjacency-masked attention.
4. Output: bounded risk weight p_{i,t} in (-1, 1) via tanh.
"""

from __future__ import annotations

import torch
from torch import nn

from .configs import ModelConfig
from .utils import causal_attention_mask


class StaticContextEncoder(nn.Module):
    """Encode per-asset static context (asset id, group id, costs)."""

    def __init__(
        self,
        *,
        n_assets: int,
        n_groups: int | None,
        cfg: ModelConfig,
    ) -> None:
        super().__init__()
        self.cfg = cfg
        self.asset_emb = nn.Embedding(n_assets, cfg.asset_embedding_dim)

        self.group_emb: nn.Embedding | None = None
        if cfg.use_group_embedding:
            if n_groups is None:
                raise ValueError("n_groups required when use_group_embedding is True")
            self.group_emb = nn.Embedding(n_groups, cfg.group_embedding_dim)

        self.include_cost = cfg.use_cost_in_context

    @property
    def context_dim(self) -> int:
        dim = self.cfg.asset_embedding_dim
        if self.group_emb is not None:
            dim += self.cfg.group_embedding_dim
        if self.include_cost:
            dim += 1
        return dim

    def forward(
        self,
        *,
        asset_ids: torch.Tensor,
        group_ids: torch.Tensor | None,
        costs: torch.Tensor | None,
    ) -> torch.Tensor:
        """Return context embedding of shape (N, C)."""
        emb_list = [self.asset_emb(asset_ids)]

        if self.group_emb is not None:
            if group_ids is None:
                raise ValueError("group_ids required when group_emb is enabled")
            emb_list.append(self.group_emb(group_ids))

        if self.include_cost:
            if costs is None:
                raise ValueError("costs required when use_cost_in_context is True")
            emb_list.append(costs)

        return torch.cat(emb_list, dim=-1)


class FiLM(nn.Module):
    """Feature-wise linear modulation: x -> x * (1 + gamma) + beta."""

    def __init__(self, *, context_dim: int, n_features: int) -> None:
        super().__init__()
        self.proj = nn.Linear(context_dim, 2 * n_features)
        self.n_features = n_features

    def forward(self, x: torch.Tensor, context: torch.Tensor) -> torch.Tensor:
        gb = self.proj(context)  # (N, 2F)
        gamma, beta = gb[:, : self.n_features], gb[:, self.n_features :]
        gamma = gamma.unsqueeze(0).unsqueeze(0)  # (1, 1, N, F)
        beta = beta.unsqueeze(0).unsqueeze(0)
        return x * (1.0 + gamma) + beta


class VectorizedVariableSelection(nn.Module):
    """Lightweight variable selection network (V-VSN)."""

    def __init__(
        self,
        *,
        n_features: int,
        d_model: int,
        context_dim: int,
        hidden_dim: int,
        dropout: float,
    ) -> None:
        super().__init__()
        self.n_features = n_features
        self.d_model = d_model

        self.feature_weight = nn.Parameter(torch.empty(n_features, d_model))
        self.feature_bias = nn.Parameter(torch.zeros(n_features, d_model))
        nn.init.xavier_uniform_(self.feature_weight)

        self.selector = nn.Sequential(
            nn.Linear(n_features + context_dim, hidden_dim),
            nn.ReLU(),
            nn.Dropout(dropout),
            nn.Linear(hidden_dim, n_features),
        )
        self.out_norm = nn.LayerNorm(d_model)

    def forward(self, x: torch.Tensor, context: torch.Tensor) -> torch.Tensor:
        b, t, n, f = x.shape
        context_bt = context.unsqueeze(0).unsqueeze(0).expand(b, t, n, -1)
        logits = self.selector(torch.cat([x, context_bt], dim=-1))
        weights = torch.softmax(logits, dim=-1)

        z = torch.einsum("btnf,fd->btnfd", x, self.feature_weight) + self.feature_bias
        h = (weights.unsqueeze(-1) * z).sum(dim=-2)
        return self.out_norm(h)


class AdapterBlock(nn.Module):
    """FFN adapter with residual connection and LayerNorm."""

    def __init__(self, *, d_model: int, hidden_mult: int, dropout: float) -> None:
        super().__init__()
        d_ff = int(hidden_mult * d_model)
        self.ln = nn.LayerNorm(d_model)
        self.ff = nn.Sequential(
            nn.Linear(d_model, d_ff), nn.GELU(), nn.Dropout(dropout), nn.Linear(d_ff, d_model)
        )
        self.dropout = nn.Dropout(dropout)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        return x + self.dropout(self.ff(self.ln(x)))


class TemporalSelfAttentionBlock(nn.Module):
    """Causal temporal self-attention per asset."""

    def __init__(self, *, d_model: int, n_heads: int, dropout: float, adapter_mult: int) -> None:
        super().__init__()
        self.mha = nn.MultiheadAttention(
            embed_dim=d_model, num_heads=n_heads, dropout=dropout, batch_first=True
        )
        self.ln = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)
        self.adapter = AdapterBlock(d_model=d_model, hidden_mult=adapter_mult, dropout=dropout)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        t = x.shape[1]
        attn_mask = causal_attention_mask(t, device=x.device)
        y, _ = self.mha(x, x, x, attn_mask=attn_mask)
        x = self.ln(x + self.dropout(y))
        return self.adapter(x)


class CrossSectionalAttention(nn.Module):
    """Cross-asset attention with Directed Delay (time lag)."""

    def __init__(self, *, d_model: int, n_heads: int, dropout: float, lag: int) -> None:
        super().__init__()
        self.lag = int(lag)
        self.mha = nn.MultiheadAttention(
            embed_dim=d_model, num_heads=n_heads, dropout=dropout, batch_first=True
        )
        self.ln = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, h: torch.Tensor, mask: torch.Tensor) -> torch.Tensor:
        b, t, n, d = h.shape

        if self.lag > 0:
            pad = torch.zeros((b, self.lag, n, d), device=h.device, dtype=h.dtype)
            h_kv = torch.cat([pad, h[:, : t - self.lag, :, :]], dim=1)
            pad_m = torch.zeros((b, self.lag, n), device=mask.device, dtype=mask.dtype)
            m_kv = torch.cat([pad_m, mask[:, : t - self.lag, :]], dim=1)
        else:
            h_kv = h
            m_kv = mask

        q = h.reshape(b * t, n, d)
        kv = h_kv.reshape(b * t, n, d)
        key_padding_mask = m_kv.reshape(b * t, n) < 0.5

        # When lag > 0, the first `lag` timesteps have all-zero keys and masks.
        # All keys masked → softmax(all -inf) → NaN. Unmask all positions for
        # those rows; attention over zero-valued keys yields zero, so the
        # residual connection passes through h unchanged.
        all_masked = key_padding_mask.all(dim=-1, keepdim=True)
        if all_masked.any():
            key_padding_mask = key_padding_mask & ~all_masked

        out, _ = self.mha(q, kv, kv, key_padding_mask=key_padding_mask)
        out = out.reshape(b, t, n, d)

        return self.ln(h + self.dropout(out))


class MacroGraphAttention(nn.Module):
    """Adjacency-masked cross-asset attention (GAT-like)."""

    def __init__(
        self,
        *,
        d_model: int,
        n_heads: int,
        dropout: float,
        adjacency_mask: torch.Tensor,
    ) -> None:
        super().__init__()
        self.register_buffer("adjacency_mask", adjacency_mask.to(dtype=torch.bool))
        self.mha = nn.MultiheadAttention(
            embed_dim=d_model, num_heads=n_heads, dropout=dropout, batch_first=True
        )
        self.ln = nn.LayerNorm(d_model)
        self.dropout = nn.Dropout(dropout)

    def forward(self, h: torch.Tensor, mask: torch.Tensor) -> torch.Tensor:
        b, t, n, d = h.shape
        x = h.reshape(b * t, n, d)
        key_padding_mask = mask.reshape(b * t, n) < 0.5

        # Guard: if adjacency_mask + key_padding_mask blocks ALL keys for any
        # query, softmax produces NaN. Unmask everything for those rows.
        combined = self.adjacency_mask.unsqueeze(0) | key_padding_mask.unsqueeze(1)
        all_blocked = combined.all(dim=-1)  # (B*T, N) — True if query i has no valid key
        if all_blocked.any():
            key_padding_mask = key_padding_mask & ~all_blocked

        out, _ = self.mha(x, x, x, attn_mask=self.adjacency_mask, key_padding_mask=key_padding_mask)
        out = out.reshape(b, t, n, d)
        return self.ln(h + self.dropout(out))


class DeepmPolicy(nn.Module):
    """DeePM policy network that outputs risk weights p_{i,t} in (-1, 1)."""

    def __init__(
        self,
        *,
        n_assets: int,
        n_features: int,
        n_groups: int | None,
        adjacency_mask: torch.Tensor | None,
        cfg: ModelConfig,
    ) -> None:
        super().__init__()
        self.n_assets = int(n_assets)
        self.n_features = int(n_features)
        self.cfg = cfg

        self.context_encoder = StaticContextEncoder(n_assets=n_assets, n_groups=n_groups, cfg=cfg)
        context_dim = self.context_encoder.context_dim

        self.film = FiLM(context_dim=context_dim, n_features=n_features)
        self.vvsn = VectorizedVariableSelection(
            n_features=n_features,
            d_model=cfg.d_model,
            context_dim=context_dim,
            hidden_dim=cfg.vvsn_hidden_dim,
            dropout=cfg.dropout,
        )

        self.lstm = nn.LSTM(
            input_size=cfg.d_model,
            hidden_size=cfg.d_model,
            num_layers=cfg.lstm_layers,
            batch_first=True,
            dropout=cfg.dropout if cfg.lstm_layers > 1 else 0.0,
        )
        self.h0_proj = nn.Linear(context_dim, cfg.lstm_layers * cfg.d_model)
        self.c0_proj = nn.Linear(context_dim, cfg.lstm_layers * cfg.d_model)

        self.temporal_blocks = nn.ModuleList(
            [
                TemporalSelfAttentionBlock(
                    d_model=cfg.d_model,
                    n_heads=cfg.n_heads,
                    dropout=cfg.dropout,
                    adapter_mult=cfg.adapter_hidden_mult,
                )
                for _ in range(cfg.temporal_mha_layers)
            ]
        )

        self.cross_attn = CrossSectionalAttention(
            d_model=cfg.d_model,
            n_heads=cfg.cross_attention_heads,
            dropout=cfg.dropout,
            lag=cfg.cross_attention_lag,
        )

        self.macro_graph: MacroGraphAttention | None = None
        if adjacency_mask is not None:
            self.macro_graph = MacroGraphAttention(
                d_model=cfg.d_model,
                n_heads=cfg.macro_gnn_heads,
                dropout=cfg.dropout,
                adjacency_mask=adjacency_mask,
            )

        self.head = nn.Linear(cfg.d_model, 1)

    def forward(
        self,
        x: torch.Tensor,
        *,
        mask: torch.Tensor,
        asset_ids: torch.Tensor,
        group_ids: torch.Tensor | None,
        costs: torch.Tensor | None,
    ) -> torch.Tensor:
        """Forward pass: features (B,T,N,F) -> risk weights (B,T,N) in (-1,1)."""
        b, t, n, f = x.shape

        context = self.context_encoder(asset_ids=asset_ids, group_ids=group_ids, costs=costs)
        x_mod = self.film(x, context)
        h = self.vvsn(x_mod, context)  # (B,T,N,D)

        # Per-asset temporal backbone
        h_bn = h.permute(0, 2, 1, 3).reshape(b * n, t, self.cfg.d_model)

        ctx_bn = context.unsqueeze(0).expand(b, n, -1).reshape(b * n, -1)
        h0 = (
            self.h0_proj(ctx_bn)
            .reshape(b * n, self.cfg.lstm_layers, self.cfg.d_model)
            .permute(1, 0, 2)
            .contiguous()
        )
        c0 = (
            self.c0_proj(ctx_bn)
            .reshape(b * n, self.cfg.lstm_layers, self.cfg.d_model)
            .permute(1, 0, 2)
            .contiguous()
        )

        h_bn, _ = self.lstm(h_bn, (h0, c0))

        for block in self.temporal_blocks:
            h_bn = block(h_bn)

        h = h_bn.reshape(b, n, t, self.cfg.d_model).permute(0, 2, 1, 3).contiguous()

        # Cross-sectional blocks
        h = self.cross_attn(h, mask)
        if self.macro_graph is not None:
            h = self.macro_graph(h, mask)

        # Output head -> tanh
        p = torch.tanh(self.head(h).squeeze(-1))
        return p * mask

```

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.