Pular para o conteúdo
Todos os documentos da biblioteca

TimeGAN para retornos financeiros sintéticos: treinamento e avaliação

Notebook Machine Learning for Trading

Resumo

Este notebook apresenta o TimeGAN, um modelo adversarial generativo recorrente para produzir sequências financeiras sintéticas. Seus cinco módulos aprendem uma representação e sua reconstrução, preveem estados latentes sucessivos, geram sequências a partir de ruído e distinguem sequências geradas das reais. O treinamento passa por aprendizado de representação, supervisão temporal e aprendizado adversarial conjunto. O exemplo usa retornos logarítmicos diários de um painel de ações dos US e reserva a parte mais recente dos dados para avaliação.

A avaliação combina verificações de distribuição com um discriminador e uma comparação de treinamento com dados sintéticos e teste com dados reais. O notebook explica por que igualar apenas médias e desvios padrão não demonstra que a estrutura serial, o agrupamento de volatilidade ou a dependência entre ações foram reproduzidos. Também compara retornos com níveis de preço: como a saída do modelo é limitada pela função sigmoide e a escala é ajustada nos dados de treinamento, níveis de preço em tendência podem colocar valores de holdout fora do intervalo do gerador e confundir a utilidade preditiva. A abordagem é apresentada como referência e tem limitações explícitas no tratamento de risco de cauda; os resultados dependem dos ativos escolhidos, da representação, das divisões e da tarefa de avaliação.

Ideias principais

  • O TimeGAN combina uma rede de incorporação, uma rede de recuperação, um supervisor, um gerador e um discriminador.
  • O treinamento separa aprendizado de representação, previsão temporal e ajuste adversarial conjunto.
  • Usam-se retornos porque a escala mínimo-máximo ajustada nos dados de treinamento pode deixar níveis de preço posteriores fora do intervalo limitado do gerador.
  • Momentos de distribuição, discriminação e utilidade preditiva posterior medem aspectos diferentes da qualidade dos dados sintéticos.
  • O método não visa especificamente reproduzir o risco de cauda.

Tags

Texto completo
# TimeGAN: Time-series Generative Adversarial Networks


# TimeGAN: Time-series Generative Adversarial Networks

**Docker image**: `ml4t-gpu`

**Book Reference**: Chapter 5, Section 5.5 (GANs for financial time series)

> **GPU recommended**: This notebook trains models with PyTorch/CUDA. It will run on CPU
> but training may be very slow. For GPU acceleration:
> ```bash
> docker compose run --rm ml4t-gpu python 05_synthetic_data/01_timegan.py
> ```


This notebook implements **TimeGAN** (Yoon, Jarrett & van der Schaar, NeurIPS 2019),
the foundational architecture for synthetic financial time series generation.

## Learning Objectives

- Understand TimeGAN's five-component architecture and why latent-space training matters
- Implement the three-phase training approach (embedding → supervisor → joint GAN)
- Evaluate synthetic data using the Fidelity-Utility-Privacy framework
- Apply Train-Synthetic-Test-Real (TSTR) validation with proper temporal splits

## Why TimeGAN Matters

TimeGAN introduced two key innovations that address limitations of standard GANs:

1. **Stepwise Supervised Loss**: Standard GANs only learn the overall distribution.
   TimeGAN adds explicit supervision on temporal transitions (how t → t+1).

2. **Learned Embedding Space**: Instead of operating directly on raw data, TimeGAN
   learns a latent representation where adversarial training is more stable.

TimeGAN remains a common baseline against which newer methods are compared.

## Data Format

We use six stocks (BA, CAT, DIS, GE, IBM, KO) and model their **daily log
returns**. The multi-stock panel exposes the model to a wider distribution of
volatility and trend regimes than single-stock OHLCV.

The choice of returns over price levels is forced by the evaluation, not a
stylistic preference. Every module here ends in a sigmoid, so the generator can
only emit values in $[0, 1]$, and the inputs are min-max scaled to that range
using statistics fitted on the training period. Price levels trend, so a later
holdout period leaves the range the scaler was fitted on. A TSTR score computed
against such a holdout is then measuring predictive utility across a severe
shift in support — the synthetic training data cannot go above one, most of the
holdout targets do — which confounds any reading of it as a statement about
temporal fidelity. Log returns are stationary enough that the holdout stays
inside the fitted range, removing the confound. Both halves of that are
measured below rather than asserted.

## References

- **Paper**: Yoon, J., Jarrett, D., & van der Schaar, M. (2019).
  "Time-series Generative Adversarial Networks." NeurIPS 2019.
- **Official Code**: https://github.com/jsyoon0823/TimeGAN

```python
"""TimeGAN — Time-series Generative Adversarial Networks."""

import hashlib
import json
from datetime import UTC, datetime
from pathlib import Path

import joblib
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
import torch
import torch.nn as nn
import torch.optim as optim
from IPython.display import Image, display
from sklearn.preprocessing import MinMaxScaler
from timegan_metrics import run_timegan_evaluation
from torch.utils.data import DataLoader, TensorDataset

from data import load_us_equities
from utils.paths import get_chapter_dir, get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, plot_fidelity_comparison, show_with_alt
```

```python
TRAIN_STEPS = 10000  # Steps per training phase (matching official repo)
RETRAIN = False  # True re-trains even when a checkpoint exists
SEED = 42
```

```python
set_global_seeds(SEED)

# Paths
ASSETS_DIR = get_chapter_dir(5) / "assets"
OUTPUT_DIR = get_output_dir(5, "timegan")
CHECKPOINT_DIR = OUTPUT_DIR / "checkpoints" / "timegan" / "multi_stock"
```

```python
# Configuration (matching 2nd edition benchmark)
#
# Data choice rationale: the six-stock panel spans different volatility/trend
# regimes, which gives the embedder a wider feature distribution than the four
# OHLCV columns of a single stock would.
TICKERS = ["BA", "CAT", "DIS", "GE", "IBM", "KO"]
SEQ_LEN = 24
HIDDEN_DIM = 24
NUM_LAYERS = 3
BATCH_SIZE = 128
LEARNING_RATE = 1e-3
D_GATING_THRESHOLD = 0.15  # Skip a discriminator step while its loss is below this

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")
```

## TimeGAN Architecture

TimeGAN consists of five interconnected modules that operate in a shared latent space:

| Module | Purpose | Training Phase |
|--------|---------|----------------|
| **Embedder** | Maps raw data → latent space | Phase 1 (autoencoder) |
| **Recovery** | Reconstructs latent → raw data | Phase 1 (autoencoder) |
| **Supervisor** | Predicts next latent step | Phase 2 (temporal) |
| **Generator** | Produces latent sequences from noise | Phase 3 (joint GAN) |
| **Discriminator** | Classifies real vs fake latent sequences | Phase 3 (joint GAN) |

```python
if (ASSETS_DIR / "timegan_architecture.jpeg").exists():
    display(Image(ASSETS_DIR / "timegan_architecture.jpeg", width=700))
```

## 1. Load Data

Six stocks with different volatility and trend regimes, converted to daily log
returns. The returned array is one row shorter than the price panel, and its
timestamps are those of the *later* day of each return.

```python
def load_price_panel(tickers: list[str], start_year: str = "2000"):
    """Adjusted close for *tickers* as a wide, gap-free panel."""
    df_pl = load_us_equities()
    df_pl = df_pl.filter(pl.col("symbol").is_in(tickers))

    # Update tickers to only those actually present in the data
    available = df_pl["symbol"].unique().to_list()
    tickers = [t for t in tickers if t in available]
    if not tickers:
        raise ValueError(
            f"None of the requested tickers found in data. Available: {available[:10]}"
        )

    # Pivot to wide format
    df = (
        df_pl.select(["timestamp", "symbol", "adj_close"])
        .pivot(on="symbol", index="timestamp", values="adj_close")
        .sort("timestamp")
        .to_pandas()
        .set_index("timestamp")
        .loc[start_year:, tickers]  # Ensure column order matches tickers
        .dropna()
    )

    return df


def load_multi_stock_data(
    tickers: list[str], start_year: str = "2000"
) -> tuple[np.ndarray, np.ndarray]:
    """Load daily log returns of adjusted close for multiple stocks."""
    df = load_price_panel(tickers, start_year)
    prices = df.values.astype(np.float32)
    assert (prices > 0).all(), "adjusted close must be positive to take logs"

    # A return is dated by the later of the two days it spans, so the first
    # price date has no return and is dropped from both arrays together.
    data = np.diff(np.log(prices), axis=0).astype(np.float32)
    timestamps = df.index.to_numpy()[1:]

    print(f"Loaded {len(df)} price rows -> {len(data)} return rows, {len(tickers)} stocks")
    print(f"Return date range: {timestamps[0]} to {timestamps[-1]}")
    print(f"Stocks: {', '.join(tickers)}")

    return data, timestamps
```

### Why Not Price Levels

Before loading returns, measure what levels would have cost. Fit the same
min-max scaler on the same training fraction of the price panel, transform the
holdout, and count the cells the generator could never reach.

```python
_levels = load_price_panel(TICKERS).values.astype(np.float32)
_split = int(len(_levels) * 0.8)
_level_holdout = MinMaxScaler().fit(_levels[:_split]).transform(_levels[_split:])
_level_seqs = np.array(
    [_level_holdout[i : i + SEQ_LEN] for i in range(len(_level_holdout) - SEQ_LEN)]
)
_usable = ((_level_seqs >= 0) & (_level_seqs <= 1)).all(axis=(1, 2))

print("If the model were trained on price levels:")
print(f"  holdout scaled range: [{_level_holdout.min():.2f}, {_level_holdout.max():.2f}]")
print(f"  holdout cells outside [0, 1]: {((_level_holdout < 0) | (_level_holdout > 1)).mean():.1%}")
print(f"  holdout sequences fully inside [0, 1]: {_usable.sum()} of {len(_level_seqs)}")
```

```python
all_data, all_timestamps = load_multi_stock_data(TICKERS)
n_features = all_data.shape[1]  # Actual number of stocks found (may differ from TICKERS)
print(f"Shape: {all_data.shape} (days × stocks)")
```

### Temporal Train/Holdout Split

We split temporally to enable unbiased TSTR evaluation.
The generator never sees holdout data during training.

```python
# Use last 20% as holdout
n_train = int(len(all_data) * 0.8)
train_data = all_data[:n_train]
holdout_data = all_data[n_train:]
train_timestamps = all_timestamps[:n_train]
holdout_timestamps = all_timestamps[n_train:]

print(f"Training:  {len(train_data):,} days ({train_timestamps[0]} to {train_timestamps[-1]})")
print(
    f"Holdout:   {len(holdout_data):,} days ({holdout_timestamps[0]} to {holdout_timestamps[-1]})"
)
```

## 2. Normalize and Create Sequences

The scaler is fitted on the training period only, so nothing about the holdout
reaches the model. That makes the holdout's scaled range a property of the
data rather than something the code can arrange, and it is the quantity the
choice of returns over levels turns on: every module ends in a sigmoid, so a
holdout cell outside $[0, 1]$ names a value the generator cannot emit and the
TSTR comparison cannot fairly ask for.

```python
scaler = MinMaxScaler()
train_scaled = scaler.fit_transform(train_data).astype(np.float32)
holdout_scaled = scaler.transform(holdout_data).astype(np.float32)

out_of_range = float(((holdout_scaled < 0) | (holdout_scaled > 1)).mean())
print(f"Train scaled range:   [{train_scaled.min():.4f}, {train_scaled.max():.4f}]")
print(f"Holdout scaled range: [{holdout_scaled.min():.4f}, {holdout_scaled.max():.4f}]")
print(f"Holdout cells outside the generator's [0, 1] range: {out_of_range:.2%}")

# A holdout that leaves the range is the failure this representation exists to
# avoid, so it stops the notebook rather than flowing into a TSTR score that
# would read as a statement about the generator.
assert out_of_range < 0.01, (
    f"{out_of_range:.1%} of holdout cells fall outside [0, 1]; a sigmoid-bounded "
    "generator cannot reach them and the TSTR comparison would measure the scaling"
)
```

```python
def create_sequences(data: np.ndarray, seq_length: int) -> np.ndarray:
    """Create overlapping sequences from time series data."""
    sequences = []
    for i in range(len(data) - seq_length):
        sequences.append(data[i : i + seq_length])
    return np.array(sequences, dtype=np.float32)
```

```python
sequences = create_sequences(train_scaled, SEQ_LEN)
holdout_sequences = create_sequences(holdout_scaled, SEQ_LEN)

print(f"Created {len(sequences)} training sequences of length {SEQ_LEN}")
print(f"Created {len(holdout_sequences)} holdout sequences for TSTR")

# DataLoader
dataset = TensorDataset(torch.from_numpy(sequences))
dataloader = DataLoader(dataset, batch_size=BATCH_SIZE, shuffle=True)
```

## 3. Model Components

We use ModuleList with separate GRU layers to match TensorFlow's layer-by-layer
construction exactly. This ensures consistent behavior with the official implementation.

```python
class Embedder(nn.Module):
    """Maps raw sequences to latent space. Stacked GRU layers + Dense with sigmoid."""

    def __init__(self, input_dim: int, hidden_dim: int, num_layers: int):
        super().__init__()
        # Stack GRU layers manually to match TF behavior exactly
        self.gru_layers = nn.ModuleList()
        for i in range(num_layers):
            in_dim = input_dim if i == 0 else hidden_dim
            self.gru_layers.append(nn.GRU(in_dim, hidden_dim, batch_first=True))
        self.fc = nn.Linear(hidden_dim, hidden_dim)

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        h = x
        for gru in self.gru_layers:
            h, _ = gru(h)
        return torch.sigmoid(self.fc(h))
```

```python
class Recovery(nn.Module):
    """Reconstructs sequences from latent space."""

    def __init__(self, hidden_dim: int, output_dim: int, num_layers: int):
        super().__init__()
        self.gru_layers = nn.ModuleList()
        for _ in range(num_layers):
            self.gru_layers.append(nn.GRU(hidden_dim, hidden_dim, batch_first=True))
        self.fc = nn.Linear(hidden_dim, output_dim)

    def forward(self, h: torch.Tensor) -> torch.Tensor:
        x = h
        for gru in self.gru_layers:
            x, _ = gru(x)
        return torch.sigmoid(self.fc(x))
```

```python
class Supervisor(nn.Module):
    """Predicts next latent step. Uses num_layers-1 layers."""

    def __init__(self, hidden_dim: int, num_layers: int):
        super().__init__()
        supervisor_layers = max(1, num_layers - 1)
        self.gru_layers = nn.ModuleList()
        for _ in range(supervisor_layers):
            self.gru_layers.append(nn.GRU(hidden_dim, hidden_dim, batch_first=True))
        self.fc = nn.Linear(hidden_dim, hidden_dim)

    def forward(self, h: torch.Tensor) -> torch.Tensor:
        s = h
        for gru in self.gru_layers:
            s, _ = gru(s)
        return torch.sigmoid(self.fc(s))
```

```python
class Generator(nn.Module):
    """Generates latent sequences from noise."""

    def __init__(self, input_dim: int, hidden_dim: int, num_layers: int):
        super().__init__()
        self.gru_layers = nn.ModuleList()
        for i in range(num_layers):
            in_dim = input_dim if i == 0 else hidden_dim
            self.gru_layers.append(nn.GRU(in_dim, hidden_dim, batch_first=True))
        self.fc = nn.Linear(hidden_dim, hidden_dim)

    def forward(self, z: torch.Tensor) -> torch.Tensor:
        e = z
        for gru in self.gru_layers:
            e, _ = gru(e)
        return torch.sigmoid(self.fc(e))
```

```python
class Discriminator(nn.Module):
    """Classifies real vs fake latent sequences."""

    def __init__(self, hidden_dim: int, num_layers: int):
        super().__init__()
        self.gru_layers = nn.ModuleList()
        for _ in range(num_layers):
            self.gru_layers.append(nn.GRU(hidden_dim, hidden_dim, batch_first=True))
        self.fc = nn.Linear(hidden_dim, 1)

    def forward(self, h: torch.Tensor) -> torch.Tensor:
        y = h
        for gru in self.gru_layers:
            y, _ = gru(y)
        return torch.sigmoid(self.fc(y))
```

## 4. Initialize Models

```python
embedder = Embedder(n_features, HIDDEN_DIM, NUM_LAYERS).to(device)
recovery = Recovery(HIDDEN_DIM, n_features, NUM_LAYERS).to(device)
supervisor = Supervisor(HIDDEN_DIM, NUM_LAYERS).to(device)
generator = Generator(n_features, HIDDEN_DIM, NUM_LAYERS).to(device)
discriminator = Discriminator(HIDDEN_DIM, NUM_LAYERS).to(device)


def count_params(model):
    return sum(p.numel() for p in model.parameters() if p.requires_grad)


print("Model Parameters:")
print(f"  Embedder:      {count_params(embedder):,}")
print(f"  Recovery:      {count_params(recovery):,}")
print(f"  Supervisor:    {count_params(supervisor):,}")
print(f"  Generator:     {count_params(generator):,}")
print(f"  Discriminator: {count_params(discriminator):,}")
total_params = sum(
    count_params(m) for m in [embedder, recovery, supervisor, generator, discriminator]
)
print(f"  Total:         {total_params:,}")
```

### Reusing a Checkpoint

A checkpoint is only interchangeable with a fresh run if it was trained on the
same thing. The weights alone cannot say what they were fitted to, so the
config saved beside them is compared field by field and a mismatch retrains
instead of loading. Without that check a checkpoint trained on price levels
would load silently into this returns-based notebook and every number below
would describe a model fitted to a different target.

The comparison has to cover everything that moves the weights, not just what
describes the data. A seed, a batch size, a learning rate and the discriminator
gating threshold all produce different models from the same inputs, so all four
are compared too. A checkpoint written before they were is missing them, which
reads as a mismatch and retrains - the safe direction.

`TICKERS` names the columns but not their contents, and the equity data is refreshed
outside this notebook. A refresh moves the training rows, which moves the scaler
fitted on them, and the run would then evaluate weights fitted under one
normalization on data scaled under another, and overwrite `scaler.pkl` with the
second one. A digest of the training rows is what closes that.

```python
CHECKPOINT_PATH = CHECKPOINT_DIR / "checkpoint.pt"
SKIP_TRAINING = False

RUN_CONFIG = {
    "tickers": TICKERS,
    "representation": "log_returns",
    "train_data_digest": hashlib.sha256(np.ascontiguousarray(train_data)).hexdigest()[:16],
    "seq_len": SEQ_LEN,
    "hidden_dim": HIDDEN_DIM,
    "num_layers": NUM_LAYERS,
    "train_steps": TRAIN_STEPS,
    # Each of these changes the weights, so a checkpoint fitted under a different
    # value is not this run's model. SEED is exposed as a papermill parameter, so
    # without it here, a second seed would load the first seed's weights.
    "seed": SEED,
    "batch_size": BATCH_SIZE,
    "learning_rate": LEARNING_RATE,
    "d_gating_threshold": D_GATING_THRESHOLD,
}

if CHECKPOINT_PATH.exists() and not RETRAIN:
    checkpoint = torch.load(CHECKPOINT_PATH, map_location=device, weights_only=False)
    saved_config = checkpoint.get("config", {})
    mismatched = {
        key: (saved_config.get(key), value)
        for key, value in RUN_CONFIG.items()
        if saved_config.get(key) != value
    }
    if "history" not in checkpoint:
        print(
            f"Checkpoint at {CHECKPOINT_PATH} predates loss-history saving; retraining so "
            "the training section has its figure."
        )
    elif mismatched:
        print(f"Checkpoint at {CHECKPOINT_PATH} does not match this run; retraining.")
        for key, (was, now) in mismatched.items():
            print(f"  {key}: checkpoint has {was!r}, this run needs {now!r}")
    else:
        print(f"Loading checkpoint from {CHECKPOINT_PATH}")
        embedder.load_state_dict(checkpoint["embedder"])
        recovery.load_state_dict(checkpoint["recovery"])
        supervisor.load_state_dict(checkpoint["supervisor"])
        generator.load_state_dict(checkpoint["generator"])
        discriminator.load_state_dict(checkpoint["discriminator"])
        history = checkpoint["history"]
        embedding_losses = history["embedding"]
        supervisor_losses = history["supervisor"]
        g_losses = history["generator"]
        d_losses = history["discriminator"]
        print("Checkpoint loaded - skipping training")
        SKIP_TRAINING = True
```

## Three-Phase Training

TimeGAN uses a sequential training approach with step-based iteration
(matching the official implementation):

1. **Embedding Phase**: Train Embedder + Recovery as autoencoder
2. **Supervisor Phase**: Train Supervisor to predict next latent step
3. **Joint Phase**: Train Generator + Discriminator adversarially

```python
if (ASSETS_DIR / "timegan_training.jpeg").exists():
    display(Image(ASSETS_DIR / "timegan_training.jpeg", width=800))
```

## 5. Phase 1: Embedding Training

Train Embedder + Recovery as autoencoder with loss: $10 \cdot \sqrt{\text{MSE}}$

```python
if not SKIP_TRAINING:
    print("\n" + "=" * 60)
    print("PHASE 1: Autoencoder Training")
    print("=" * 60)

    opt_autoencoder = optim.Adam(
        list(embedder.parameters()) + list(recovery.parameters()),
        lr=LEARNING_RATE,
    )
    mse_loss = nn.MSELoss()

    # Infinite iterator for step-based training
    def infinite_dataloader():
        while True:
            yield from dataloader

    data_iter = infinite_dataloader()
    embedding_losses = []

    for step in range(TRAIN_STEPS):
        (batch,) = next(data_iter)
        batch = batch.to(device)

        h = embedder(batch)
        x_tilde = recovery(h)

        # Loss: 10 * sqrt(MSE) - matching 2nd edition exactly
        embedding_loss = mse_loss(x_tilde, batch)
        e_loss = 10.0 * torch.sqrt(embedding_loss)

        opt_autoencoder.zero_grad()
        e_loss.backward()
        opt_autoencoder.step()

        if step % 1000 == 0:
            embedding_losses.append(torch.sqrt(embedding_loss).item())
            print(f"  Step {step}: loss = {embedding_losses[-1]:.6f}", flush=True)

    print(f"Phase 1 complete. Final loss: {embedding_losses[-1]:.6f}")
```

## 6. Phase 2: Supervisor Training

Train Supervisor to predict next latent step: $\mathcal{L}_S = ||h_{t+1} - \hat{s}_t||_2$

```python
if not SKIP_TRAINING:
    print("\n" + "=" * 60)
    print("PHASE 2: Supervisor Training")
    print("=" * 60)

    opt_supervisor = optim.Adam(supervisor.parameters(), lr=LEARNING_RATE)
    supervisor_losses = []

    for step in range(TRAIN_STEPS):
        (batch,) = next(data_iter)
        batch = batch.to(device)

        with torch.no_grad():
            h = embedder(batch)

        h_hat = supervisor(h)
        g_loss_s = mse_loss(h[:, 1:, :], h_hat[:, :-1, :])

        opt_supervisor.zero_grad()
        g_loss_s.backward()
        opt_supervisor.step()

        if step % 1000 == 0:
            supervisor_losses.append(g_loss_s.item())
            print(f"  Step {step}: loss = {g_loss_s.item():.6f}", flush=True)

    print(f"Phase 2 complete. Final loss: {supervisor_losses[-1]:.6f}")
```

## 7. Phase 3: Joint Adversarial Training

Train Generator and Discriminator adversarially while maintaining reconstruction
and supervised losses. Key details:
- 2 generator+embedder updates per discriminator update

  exceeds `D_GATING_THRESHOLD`, which keeps an already-winning
  discriminator from overwhelming the generator

```python
def get_moment_loss(y_true: torch.Tensor, y_pred: torch.Tensor) -> torch.Tensor:
    """Match first two moments between real and synthetic."""
    y_true_mean = y_true.mean(dim=0)
    y_pred_mean = y_pred.mean(dim=0)
    y_true_var = y_true.var(dim=0)
    y_pred_var = y_pred.var(dim=0)

    loss_mean = torch.abs(y_true_mean - y_pred_mean).mean()
    loss_var = torch.abs(torch.sqrt(y_true_var + 1e-6) - torch.sqrt(y_pred_var + 1e-6)).mean()

    return loss_mean + loss_var
```

```python
if not SKIP_TRAINING:
    print("\n" + "=" * 60)
    print("PHASE 3: Joint Adversarial Training")
    print("=" * 60)

    opt_generator = optim.Adam(
        list(generator.parameters()) + list(supervisor.parameters()),
        lr=LEARNING_RATE,
    )
    opt_discriminator = optim.Adam(discriminator.parameters(), lr=LEARNING_RATE)
    opt_embedder = optim.Adam(
        list(embedder.parameters()) + list(recovery.parameters()),
        lr=LEARNING_RATE,
    )

    bce_loss = nn.BCELoss()
    gamma = 1.0
    g_losses, d_losses = [], []
```

```python
if not SKIP_TRAINING:
    for step in range(TRAIN_STEPS):
        # 2 generator+embedder updates per discriminator update
        for _ in range(2):
            (batch,) = next(data_iter)
            batch = batch.to(device)
            batch_size_actual = batch.shape[0]

            z = torch.rand(batch_size_actual, SEQ_LEN, n_features, device=device)

            # Generator step
            opt_generator.zero_grad()

            # Generate synthetic latent sequences
            e_hat = generator(z)
            h_hat = supervisor(e_hat)

            # Supervised loss on SYNTHETIC: generator learns temporal coherence
            # Predict e_hat[t+1] from e_hat[t], so align e_hat[1:] with h_hat[:-1]
            g_loss_s = mse_loss(e_hat[:, 1:, :], h_hat[:, :-1, :])

            y_fake = discriminator(h_hat)
            y_fake_e = discriminator(e_hat)

            g_loss_u = bce_loss(y_fake, torch.ones_like(y_fake))
            g_loss_u_e = bce_loss(y_fake_e, torch.ones_like(y_fake_e))

            x_hat = recovery(h_hat)
            g_loss_v = get_moment_loss(batch, x_hat)

            g_loss = g_loss_u + g_loss_u_e + 100.0 * torch.sqrt(g_loss_s) + 100.0 * g_loss_v
            g_loss.backward()
            opt_generator.step()

            # Embedder step
            opt_embedder.zero_grad()

            h = embedder(batch)
            h_hat_sup = supervisor(h)
            # Supervised loss: predict h[t+1] from h[t], so align h[1:] with h_hat_sup[:-1]
            g_loss_s = mse_loss(h[:, 1:, :], h_hat_sup[:, :-1, :])

            x_tilde = recovery(h)
            e_loss_t0 = mse_loss(x_tilde, batch)

            e_loss = 10.0 * torch.sqrt(e_loss_t0) + 0.1 * g_loss_s
            e_loss.backward()
            opt_embedder.step()

        # Discriminator step (with gating)
        (batch,) = next(data_iter)
        batch = batch.to(device)
        batch_size_actual = batch.shape[0]
        z = torch.rand(batch_size_actual, SEQ_LEN, n_features, device=device)

        with torch.no_grad():
            h = embedder(batch)
            e_hat = generator(z)
            h_hat = supervisor(e_hat)

        y_real = discriminator(h)
        y_fake = discriminator(h_hat)
        y_fake_e = discriminator(e_hat)

        d_loss_real = bce_loss(y_real, torch.ones_like(y_real))
        d_loss_fake = bce_loss(y_fake, torch.zeros_like(y_fake))
        d_loss_fake_e = bce_loss(y_fake_e, torch.zeros_like(y_fake_e))

        d_loss = d_loss_real + d_loss_fake + gamma * d_loss_fake_e

        if d_loss.item() > D_GATING_THRESHOLD:
            opt_discriminator.zero_grad()
            d_loss.backward()
            opt_discriminator.step()

        if step % 1000 == 0:
            g_losses.append(g_loss.item())
            d_losses.append(d_loss.item())
            print(
                f"  Step {step:6,}: D={d_loss.item():.4f}, G={g_loss.item():.4f}, "
                f"g_s={g_loss_s.item():.4f}, g_v={g_loss_v.item():.4f}",
                flush=True,
            )

    print(f"Phase 3 complete. Final D={d_losses[-1]:.4f}, G={g_losses[-1]:.4f}")
```

## 8. Generate Synthetic Data

Training draws millions of random numbers and loading a checkpoint draws none, so the
two paths arrive here with different RNG states. Everything below - the noise the
generator is fed, and the initializations inside the evaluation suite - would then
differ between a run that trained the weights and a run that loaded the same weights
back. Reseeding here makes the two paths produce the same synthetic sequences.

```python
set_global_seeds(SEED)

print("\n=== Generating Synthetic Data ===")

generator.eval()
supervisor.eval()
recovery.eval()

n_synthetic = len(sequences)
generated_batches = []

with torch.no_grad():
    for i in range(0, n_synthetic, BATCH_SIZE):
        batch_size = min(BATCH_SIZE, n_synthetic - i)
        z = torch.rand(batch_size, SEQ_LEN, n_features, device=device)
        e_hat = generator(z)
        h_hat = supervisor(e_hat)
        x_hat = recovery(h_hat)
        generated_batches.append(x_hat.cpu().numpy())

synthetic = np.vstack(generated_batches)
print(f"Generated {len(synthetic)} synthetic sequences")
print(f"Synthetic mean: {synthetic.mean():.4f} (real: {sequences.mean():.4f})")
print(f"Synthetic std:  {synthetic.std():.4f} (real: {sequences.std():.4f})")
```

## 9. Evaluation

We evaluate using the Fidelity-Utility-Privacy framework with LSTM-based
evaluation matching the original paper.

### Diversity: PCA and t-SNE Projections

```python
fig = plot_fidelity_comparison(
    sequences, synthetic, title="Real and synthetic sequences in two projections", n_samples=1000
)
show_with_alt(
    fig,
    "Two scatter panels of the same real and synthetic sequences. In the PCA "
    "projection the real points form a broad cloud around the origin while the "
    "synthetic points sit in a short horizontal sliver at its centre. In the t-SNE "
    "projection the two occupy separate regions of the plane, synthetic to the left "
    "and real to the right, with almost no interleaving.",
)
```

**Read it.** Neither panel shows the two sets sitting on top of each other. In the
PCA projection the synthetic sequences occupy a sliver at the centre of the real
cloud, so they vary far less than the real ones along the directions that carry most
of the real variance. In the t-SNE projection the two sets fall in separate regions.
The discriminative accuracy printed by the evaluation suite below puts a number on
what these panels show.

Read this against the mean and standard deviation printed by the generation cell
above, which are close to the real ones. Those two readings are not in conflict, and
the section at the end of the notebook is about why: the pooled moments are a weak
requirement, and these panels are a picture of what they leave unconstrained.

### Paper Evaluation Suite (LSTM-based)

Run the full evaluation following Yoon et al. (2019) using LSTM predictors
and discriminators, not tree-based models.

```python
print("=" * 70)
print("TIMEGAN EVALUATION SUITE")
print("=" * 70)

# Evaluation settings per EXPERIMENTS.md:
# - hidden_dim=input_dim (small classifier to avoid overfitting)
# - epochs_disc=250 (matches benchmark)
eval_results = run_timegan_evaluation(
    synthetic=synthetic,
    real_train=sequences,
    real_holdout=holdout_sequences,
    hidden_dim=n_features,  # Must match input_dim for fair evaluation
    epochs_disc=250,  # Per benchmark experiments
    quick_test=(TRAIN_STEPS < 1000),
    verbose=True,
    include_yoon=True,
)

disc_accuracy = eval_results["discriminative"]["accuracy"]
tstr_ratio = eval_results["predictive"]["ratio"]

print("\n" + "=" * 70)
print("SUMMARY")
print("=" * 70)
print(f"Discriminative Accuracy: {disc_accuracy:.1%} (target: ~50%)")
print(f"TSTR Ratio: {tstr_ratio:.3f} (target: ~1.0)")
```

### Training Curves

```python
if embedding_losses:
    fig, axes = plt.subplots(1, 3, figsize=(14, 4))

    axes[0].plot(embedding_losses, color=COLORS["blue"], linewidth=1.5)
    axes[0].set_title("Phase 1: Embedding")
    axes[0].set_xlabel("Step (×1000)")
    axes[0].set_ylabel("Loss")

    axes[1].plot(supervisor_losses, color=COLORS["blue"], linewidth=1.5)
    axes[1].set_title("Phase 2: Supervisor")
    axes[1].set_xlabel("Step (×1000)")

    axes[2].plot(g_losses, label="Generator", color=COLORS["blue"], linewidth=1.5)
    axes[2].plot(d_losses, label="Discriminator", color=COLORS["amber"], linewidth=1.5)
    axes[2].set_title("Phase 3: Joint")
    axes[2].set_xlabel("Step (×1000)")
    axes[2].legend()

    fig.suptitle("Loss curves for the three training phases", fontsize=14, fontweight="semibold")
    show_with_alt(
        fig,
        "Three panels, each against training step in thousands and each on its own "
        "vertical scale. The embedding autoencoder loss declines unevenly across the "
        "whole run; the supervisor loss drops sharply inside the first thousand steps "
        "and is flat after that; the third panel plots the generator and discriminator "
        "losses of the joint phase together, the generator falling steeply inside the "
        "first thousand steps and then running flat, and the discriminator flat across "
        "the whole run at a level a little below where the generator settles.",
    )
```

## 10. Save Outputs

```python
CHECKPOINT_DIR.mkdir(parents=True, exist_ok=True)

# Save checkpoint
checkpoint = {
    "embedder": embedder.state_dict(),
    "recovery": recovery.state_dict(),
    "supervisor": supervisor.state_dict(),
    "generator": generator.state_dict(),
    "discriminator": discriminator.state_dict(),
    "config": RUN_CONFIG,
    # The training section plots these. Without them a checkpoint-loading run renders
    # that section with no figure, so they travel with the weights they describe.
    "history": {
        "embedding": embedding_losses,
        "supervisor": supervisor_losses,
        "generator": g_losses,
        "discriminator": d_losses,
    },
}
torch.save(checkpoint, CHECKPOINT_PATH)

# Save scaler for denormalization
joblib.dump(scaler, CHECKPOINT_DIR / "scaler.pkl")
```

```python
# Save metadata
metadata = {
    "generator": "timegan",
    "paper": "Yoon et al., NeurIPS 2019",
    "created_at": datetime.now(UTC).isoformat(),
    "data_format": "multi_stock_adj_close_log_returns",
    "tickers": TICKERS,
    "config": RUN_CONFIG,
    "evaluation": {
        "discriminative_accuracy": float(disc_accuracy),
        "tstr_ratio": float(tstr_ratio),
    },
}
with open(CHECKPOINT_DIR / "metadata.json", "w") as f:
    json.dump(metadata, f, indent=2)

# Save samples
np.save(CHECKPOINT_DIR / "synthetic.npy", synthetic)
np.save(CHECKPOINT_DIR / "real_train.npy", sequences)
np.save(CHECKPOINT_DIR / "real_holdout.npy", holdout_sequences)
```

### Persisting the Projection Coordinates

The book-repo Hard Rule 15 script
`05_synthetic_data/figures/scripts/generate_figure_5_04_timegan_fidelity.py`
re-renders the publication figure without retraining. It reads vendored copies
under `05_synthetic_data/figures/data/timegan_fidelity/`, so it does not track
these arrays until someone re-vendors them.

The projections are computed with the parameters `plot_fidelity_comparison()`
uses internally, so the inline render and the persisted arrays describe the
same projection. That includes the legacy global-RNG seeding, `np.random.seed`
followed by `np.random.choice`, so the subsample indices follow the helper's
MT19937 sequence rather than the PCG64 sequence `np.random.default_rng` emits.

```python
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE

_n_viz = min(1000, len(sequences), len(synthetic))
np.random.seed(42)
_idx_real = np.random.choice(len(sequences), _n_viz, replace=False)
_idx_synth = np.random.choice(len(synthetic), _n_viz, replace=False)
_real_flat = sequences[_idx_real].mean(axis=1)
_synth_flat = synthetic[_idx_synth].mean(axis=1)

_pca = PCA(n_components=min(2, _real_flat.shape[1], _n_viz))
_pca.fit(_real_flat)
_real_pca = _pca.transform(_real_flat)
_synth_pca = _pca.transform(_synth_flat)

_combined = np.vstack([_real_flat, _synth_flat])
_tsne = TSNE(
    n_components=min(2, _real_flat.shape[1]),
    perplexity=min(40, max(2, _n_viz // 4)),
    max_iter=1000,
    random_state=42,
)
_combined_tsne = _tsne.fit_transform(_combined)
_real_tsne = _combined_tsne[:_n_viz]
_synth_tsne = _combined_tsne[_n_viz:]

np.save(CHECKPOINT_DIR / "fidelity_real_pca.npy", _real_pca)
np.save(CHECKPOINT_DIR / "fidelity_synth_pca.npy", _synth_pca)
np.save(CHECKPOINT_DIR / "fidelity_real_tsne.npy", _real_tsne)
np.save(CHECKPOINT_DIR / "fidelity_synth_tsne.npy", _synth_tsne)

print(f"\nSaved to {CHECKPOINT_DIR}/")
```

## Summary

This notebook implemented **TimeGAN** (Yoon et al., NeurIPS 2019):

1. **Data**: 6 diverse stocks (BA, CAT, DIS, GE, IBM, KO) with adjusted close
2. **Architecture**: Five-component system with ModuleList GRUs (matching TF)
3. **Training**: Three-phase, step-based approach (10,000 steps per phase)
4. **Evaluation**: LSTM-based discriminative and predictive scores

### Reading the Two Scores

Both diagnostics are printed above rather than quoted here, because both move
from run to run and a number typed into prose stops tracking the code that
produced it. What is fixed is how to read them.

**Discriminative accuracy** trains a classifier to tell real sequences from
synthetic ones and reports its held-out accuracy. The benchmark is chance, one
half, not zero: a value near one half means the classifier cannot separate the
two, and higher means it can. Held-out accuracy is not bounded below by chance,
so a value well under one half is possible and is not a better result — it
points at sampling variation, a classifier that failed to generalize, or a
shift between the two sets. A substantial departure in either direction is
something to investigate rather than to report.

**The TSTR/TRTR ratio** divides the error of a predictor trained on synthetic
data by the error of the same predictor trained on real data, both scored on
the real holdout. Its target is one, meaning synthetic data substitutes for
real without loss; above one means it is worse.

The predictor's output layer is linear and unbounded, so it can predict outside
the generator's range and the ratio does not become meaningless when the
holdout leaves it. What it becomes is confounded. On price levels the synthetic
training data is confined to $[0, 1]$ while most holdout targets sit above one,
as the cell before the data load shows, and a ratio measured across that gap
cannot separate a generator that missed the temporal structure from targets
that left the support the synthetic data covers. Working in returns keeps the
holdout inside the fitted range, which the assertion in the normalization
section enforces, so the ratio is answering one question rather than two.

### When the Moments Agree and the Discriminator Does Not

The generation cell prints the synthetic and real mean and standard deviation,
and on this run they agree to about two decimal places. The discriminative
accuracy printed just above is nonetheless far from chance. Those two facts are
not in conflict, and holding them together is the point of the section the
chapter devotes to evaluation.

Matching the first two moments of a pooled distribution is a weak requirement.
It says nothing about the order of values within a sequence, about how
volatility clusters, or about the dependence between the six stocks on the same
day. A classifier that reads whole sequences can use any of that, and a high
accuracy says it found something it could use. Which of those the generator
missed is not something the accuracy alone can say — it is a single number, and
diagnosing it needs the stylized-fact and dependence checks the chapter
introduces alongside this score.

This is also why the two diagnostics are reported together rather than one
being chosen. A ratio near one says a downstream predictor trained on synthetic
data transfers to real data for this particular task, which is a claim about
utility. It does not say the synthetic sequences are indistinguishable from
real ones, which is what the discriminator tests.

### Limitations

TimeGAN focuses on matching overall distribution, not tail risk. For alternatives:
- **Tail risk**: See [`02_tailgan_tail_risk`](02_tailgan_tail_risk.ipynb)
- **Path signatures**: See [`03_sigcwgan_signatures`](03_sigcwgan_signatures.ipynb)
- **Diffusion models**: See [`05_diffusion_ts`](05_diffusion_ts.ipynb)
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.