OLS诊断与稳健推断:ETF收益面板
代码 《交易机器学习》
总结
本笔记对 ETF 特征和前向收益应用经典普通最小二乘推断。它回顾系数估计、标准误、t 统计量、p 值,以及对异方差、残差自相关、非正态性和多重共线性的检验。样本是一个不平衡的资产日期面板;工作流程采用滚动训练折,使用训练数据标准化特征,并在估计稳健协方差指标前剔除若干线性相关的收益合成指标。
讨论强调,诊断结果取决于面板观测值的排序和分组方式。它说明了重叠的前向收益标签如何导致相邻日期的残差相关,并解释为何在合并后的行序上应用常见的序列相关检验或协方差修正可能产生误导。在非球形误差下,稳健标准误可以改善不确定性估计,但无法修复内生性、函数形式设定错误或较弱的预测表现。本笔记的证据仅适用于一个 ETF 案例研究和一个数据折;推断检验不能证明样本外排名能力,这需要另行进行预测评估。
核心观点
- OLS 推断概括系数的不确定性,但不衡量预测质量。
- 定义时间顺序和横截面相关性时,不平衡的 ETF 面板需要谨慎处理。
- 重叠的前向收益标签可能导致相邻观测之间出现残差自相关。
- 稳健标准误可以处理某些协方差问题,但无法修复不一致估计或改善预测。
- 多重共线性诊断可以发现冗余或高度相关的特征。
标签
全文
# 07_dp_gan.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# formats: ipynb,py:percent
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Chapter 5: Differential Privacy for Generative Models
#
# **Chapter 5: Synthetic Data Generation**
# **Section Reference**: Section 5.8 (Applying the Fidelity-Utility-Privacy framework)
#
# **Docker image**: `ml4t-gpu`
#
# > **GPU recommended**: This notebook trains models with PyTorch/CUDA. It will run on CPU
# > but training may be very slow. For GPU acceleration:
# > ```bash
# > docker compose run --rm ml4t-gpu python 05_synthetic_data/07_dp_gan.py
# > ```
#
#
# ## Purpose
#
# This notebook demonstrates **Differential Privacy (DP)** for training generative models
# using **Opacus**, PyTorch's official DP library. We train a GAN with formal privacy
# guarantees that limit information leakage about individual records.
#
# ## Learning Objectives
#
# By completing this notebook, you will:
# - Understand the differential privacy framework and privacy budget (ε)
# - Implement DP-SGD training using Opacus (proper per-sample gradients)
# - Observe the privacy-utility tradeoff as ε varies
# - Evaluate synthetic data quality under privacy constraints
#
# ## Cross-References
#
# - **Upstream**: ETF Universe loader (`data`)
# - **Book**: Section 5.8 discusses differential privacy in the FUP framework
#
# ---
#
# ## Key Concepts
#
# 1. **Differential Privacy**: Mathematical framework for privacy guarantees
# 2. **DP-SGD**: Per-sample gradient clipping + noise (NOT batch-level)
# 3. **Privacy Budget (ε)**: Trade-off between privacy and utility
# 4. **Privacy Accountant**: Tracks cumulative privacy cost during training
#
# ## Why Opacus?
#
# A correct DP-SGD implementation requires:
# - **Per-sample gradients**: Not available in standard PyTorch
# - **Proper clipping**: Before aggregation, not after
# - **Privacy accounting**: Accurate ε computation via moments accountant
#
# Opacus handles all of this correctly. Manual implementations are error-prone.
#
# ## References
#
# - Abadi et al. (2016). "Deep Learning with Differential Privacy"
# - [Opacus Documentation](https://opacus.ai/)
# - Yousefpour et al. (2021). "Opacus: User-Friendly Differential Privacy Library in PyTorch"
# %%
"""Differential Privacy for Generative Models — DP-GAN with Opacus privacy guarantees."""
import logging
import warnings
from datetime import date
# Scoped by category and module, so a warning from this notebook's own code still shows.
warnings.filterwarnings("ignore", category=UserWarning, module="opacus")
warnings.filterwarnings("ignore", category=UserWarning, module="torch")
# The hook warning is raised from C, so it is attributed to `sys` rather than to torch and
# the module filter above never sees it. It fires because opacus registers backward hooks on
# the discriminator's first layer, whose input is data and requires no gradient.
warnings.filterwarnings("ignore", message="Full backward hook is firing")
# Opacus reports the drop_last it ignores through its own logger, not through warnings.
# DPDataLoader samples with Poisson sampling, so batch size varies and drop_last has no meaning.
logging.getLogger("opacus").setLevel(logging.ERROR)
import matplotlib.pyplot as plt
import numpy as np
import plotly.graph_objects as go
import polars as pl
import torch
import torch.nn as nn
from opacus import PrivacyEngine
from opacus.validators import ModuleValidator
from plotly.subplots import make_subplots
from torch.utils.data import DataLoader, TensorDataset
from tqdm import tqdm
from data import load_etfs
from utils.paths import get_output_dir
from utils.reproducibility import set_global_seeds
from utils.style import COLORS, plot_fidelity_comparison, show_plotly_with_alt, show_with_alt
# %% tags=["parameters"]
# DP-GAN parameters (Abadi et al. 2016)
HIDDEN_DIM = 128 # Generator/discriminator hidden dimension
LATENT_DIM = 32 # Noise vector dimension
EPOCHS = 20 # Training epochs
# 0 = use default list; >0 limits symbol count
MAX_SYMBOLS = 0
SEED = 42
# Progress bars write to stderr and papermill records every repaint. The training loop
# prints its losses and spent budget every few epochs, so nothing is lost with them off.
PROGRESS_BARS = False
# %%
set_global_seeds(SEED)
# %%
# Configuration
CONFIG = {
# Data
"start_date": "2020-01-01",
"n_features": 6, # Number of features to use
# Model
"hidden_dim": HIDDEN_DIM,
"latent_dim": LATENT_DIM,
# Training
"epochs": EPOCHS,
"batch_size": 64,
"learning_rate": 1e-3,
# Privacy parameters
"epsilon": 10.0, # Target privacy budget
"delta": 1e-5, # Probability of privacy breach
"max_grad_norm": 1.0, # Per-sample gradient clipping threshold
}
# Checkpoint configuration
RETRAIN = False # Set True to retrain even if checkpoint exists
CHECKPOINT_PATH = get_output_dir(5, "dp_gan") / "checkpoints" / "dp_gan_model.pt"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")
print(f"Target privacy: (ε={CONFIG['epsilon']}, δ={CONFIG['delta']})")
# %% [markdown]
# ## 1. Differential Privacy Background
#
# A randomized algorithm $\mathcal{M}$ is $(\varepsilon, \delta)$-differentially private
# if for any two adjacent datasets $D$ and $D'$ (differing by one record) and any output $S$:
#
# $$P[\mathcal{M}(D) \in S] \leq e^\varepsilon \cdot P[\mathcal{M}(D') \in S] + \delta$$
#
# - **ε (epsilon)**: Privacy budget. Lower = more privacy, less utility.
# - **δ (delta)**: Probability of privacy breach. Should be << 1/n.
#
# ### Common ε Values
#
# | ε Range | Privacy Level | Use Case |
# |---------|---------------|----------|
# | ε < 1 | Very Strong | Highly sensitive data (medical, financial PII) |
# | 1 ≤ ε ≤ 10 | Moderate | Most production applications |
# | ε > 10 | Weak | When utility is paramount |
#
# ### Why DP-SGD Requires Per-Sample Gradients
#
# Standard gradient clipping clips the **sum** of gradients. This doesn't bound
# the contribution of any individual sample. DP-SGD clips **each sample's gradient
# independently**, then adds noise proportional to the clipping threshold.
# %% [markdown]
# ## 2. Load Real Financial Data
# %%
# Load ETF data
print("Loading ETF data...")
df = load_etfs()
start_dt = date.fromisoformat(CONFIG["start_date"])
# Filter to recent data and high-volume ETFs
_DEFAULT_SYMBOLS = ["SPY", "QQQ", "IWM", "EFA", "TLT", "GLD", "XLF", "XLE", "XLK", "XLV"]
SYMBOLS = _DEFAULT_SYMBOLS[:MAX_SYMBOLS] if MAX_SYMBOLS > 0 else _DEFAULT_SYMBOLS
df = df.filter((pl.col("timestamp") >= start_dt) & pl.col("symbol").is_in(SYMBOLS))
print(f"Loaded {len(df)} rows for {len(SYMBOLS)} ETFs")
# %% [markdown]
# ## 3. Engineer Features for GAN Training
# %%
# Create features per symbol
features_list = []
for symbol in SYMBOLS:
sym_df = df.filter(pl.col("symbol") == symbol).sort("timestamp")
if len(sym_df) < 100:
continue
# Compute returns and indicators
sym_df = sym_df.with_columns(
[
# Returns
(pl.col("close").pct_change() * 100).alias("ret_1d"),
(pl.col("close").pct_change(5) * 100).alias("ret_5d"),
# Volatility (20-day rolling std of returns)
(pl.col("close").pct_change().rolling_std(20) * 100 * np.sqrt(252)).alias("volatility"),
# Volume ratio
(pl.col("volume") / pl.col("volume").rolling_mean(20)).alias("volume_ratio"),
# Price momentum (normalized)
(
(pl.col("close") - pl.col("close").rolling_mean(20))
/ pl.col("close").rolling_std(20)
).alias("momentum_z"),
# High-low range (normalized)
((pl.col("high") - pl.col("low")) / pl.col("close") * 100).alias("range_pct"),
]
)
# Drop nulls and select features
feature_cols = ["ret_1d", "ret_5d", "volatility", "volume_ratio", "momentum_z", "range_pct"]
sym_features = sym_df.select(feature_cols).drop_nulls()
features_list.append(sym_features)
# Combine all symbols
all_features = pl.concat(features_list)
print(f"Total feature samples: {len(all_features)}")
# %%
# Convert to numpy
feature_names = all_features.columns
real_data = all_features.to_numpy().astype(np.float32)
# Handle any remaining infinities
real_data = np.nan_to_num(real_data, nan=0.0, posinf=3.0, neginf=-3.0)
# Normalize for GAN training
data_mean = real_data.mean(axis=0)
data_std = real_data.std(axis=0) + 1e-8
data_normalized = (real_data - data_mean) / data_std
# Clip outliers
data_normalized = np.clip(data_normalized, -3, 3)
print(f"Feature matrix shape: {data_normalized.shape}")
print(f"Features: {feature_names}")
# Create DataLoader
dataset = TensorDataset(torch.FloatTensor(data_normalized))
# Note: Opacus works best with fixed batch sizes
train_loader = DataLoader(dataset, batch_size=CONFIG["batch_size"], shuffle=True, drop_last=True)
# %% [markdown]
# ## 4. GAN Architecture (Opacus-Compatible)
#
# Opacus has specific requirements for model architecture:
# - No in-place operations
# - Use `nn.GroupNorm` instead of `nn.BatchNorm` (BatchNorm breaks DP)
# - Models must pass `ModuleValidator.validate()`
# %% [markdown]
# ### Generator
#
# The generator maps random noise to synthetic feature vectors. It uses `Tanh` output
# scaled to [-3, 3] to match our normalized data range. The generator does not need
# DP modifications because it never sees real data directly.
# %%
class Generator(nn.Module):
"""Generator network for tabular data."""
def __init__(self, latent_dim: int, hidden_dim: int, output_dim: int):
super().__init__()
self.net = nn.Sequential(
nn.Linear(latent_dim, hidden_dim),
nn.LeakyReLU(0.2),
nn.Linear(hidden_dim, hidden_dim),
nn.LeakyReLU(0.2),
nn.Linear(hidden_dim, output_dim),
nn.Tanh(), # Output in [-1, 1] range
)
def forward(self, z: torch.Tensor) -> torch.Tensor:
return self.net(z) * 3 # Scale to [-3, 3] to match normalized data
# %% [markdown]
# ### Discriminator (DP-Compatible)
#
# The discriminator must be DP-compatible because it processes real data. The key
# architectural constraint: **BatchNorm is replaced with GroupNorm**. BatchNorm computes
# statistics across the batch, which leaks information about other samples and violates
# the per-sample privacy guarantee. GroupNorm operates within each sample independently.
# %%
class Discriminator(nn.Module):
"""
Discriminator network.
Note: We use GroupNorm instead of BatchNorm because BatchNorm
violates differential privacy (it leaks information about other samples).
"""
def __init__(self, input_dim: int, hidden_dim: int):
super().__init__()
self.net = nn.Sequential(
nn.Linear(input_dim, hidden_dim),
nn.GroupNorm(1, hidden_dim), # DP-safe normalization
nn.LeakyReLU(0.2),
nn.Linear(hidden_dim, hidden_dim),
nn.GroupNorm(1, hidden_dim),
nn.LeakyReLU(0.2),
nn.Linear(hidden_dim, 1),
)
def forward(self, x: torch.Tensor) -> torch.Tensor:
return self.net(x)
# %% [markdown]
# ### Initialize and Validate Models
#
# Opacus provides `ModuleValidator` to check whether a model is compatible with
# DP-SGD and automatically fix common issues (e.g., replacing any remaining
# BatchNorm layers).
# %%
# Initialize models
input_dim = data_normalized.shape[1]
generator = Generator(CONFIG["latent_dim"], CONFIG["hidden_dim"], input_dim).to(device)
discriminator = Discriminator(input_dim, CONFIG["hidden_dim"]).to(device)
# Validate discriminator is compatible with Opacus
errors = ModuleValidator.validate(discriminator, strict=False)
if errors:
print("Opacus compatibility issues:")
for e in errors:
print(f" - {e}")
# Fix automatically
discriminator = ModuleValidator.fix(discriminator)
print("Fixed automatically.")
else:
print("Discriminator is Opacus-compatible")
discriminator = discriminator.to(device)
# %% [markdown]
# ## 5. Train with DP-SGD using Opacus
#
# Opacus wraps the model, optimizer, and dataloader to:
# 1. Compute per-sample gradients via `functorch`
# 2. Clip each sample's gradient to `max_grad_norm`
# 3. Add calibrated Gaussian noise
# 4. Track privacy budget via moments accountant
#
# The Lipschitz constraint the WGAN objective needs is enforced by clipping the
# discriminator's weights rather than by a gradient penalty. That is a privacy decision,
# not a modelling preference: a gradient penalty differentiates through interpolations
# between real and generated samples, which puts real records into a term Opacus is not
# accounting for. Weight clipping touches only the parameters, so the privacy accounting
# stays sound.
# %%
def train_dp_gan(
generator: nn.Module,
discriminator: nn.Module,
train_loader: DataLoader,
epochs: int,
target_epsilon: float,
target_delta: float,
max_grad_norm: float,
) -> dict:
"""
Train GAN with DP-SGD using Opacus.
The discriminator is trained with differential privacy (it sees real data).
The generator is trained using a separate discriminator copy to avoid
triggering Opacus hooks during backprop.
This is the standard approach for DP-GANs:
1. Train discriminator with DP on real/fake discrimination
2. Copy weights to a clean discriminator for generator training
3. Generator backprops through clean copy (no privacy impact)
"""
# Create a clean discriminator copy for generator training (no DP hooks)
discriminator_for_g = Discriminator(input_dim, CONFIG["hidden_dim"]).to(device)
# Optimizers
opt_g = torch.optim.Adam(generator.parameters(), lr=CONFIG["learning_rate"], betas=(0.5, 0.999))
opt_d = torch.optim.Adam(
discriminator.parameters(), lr=CONFIG["learning_rate"], betas=(0.5, 0.999)
)
# Wrap discriminator with Opacus PrivacyEngine
# Use RDP accountant for numerical stability with small datasets
privacy_engine = PrivacyEngine(accountant="rdp")
discriminator_private, opt_d, train_loader_private = privacy_engine.make_private_with_epsilon(
module=discriminator,
optimizer=opt_d,
data_loader=train_loader,
target_epsilon=target_epsilon,
target_delta=target_delta,
epochs=epochs,
max_grad_norm=max_grad_norm,
)
print("\nOpacus privacy parameters:")
print(f" Noise multiplier: {opt_d.noise_multiplier:.4f}")
print(f" Max grad norm: {max_grad_norm}")
print(f" Target: (ε={target_epsilon}, δ={target_delta})")
history = {"g_loss": [], "d_loss": [], "epsilon": []}
for epoch in tqdm(range(epochs), desc="Training DP-GAN", disable=not PROGRESS_BARS):
epoch_g_loss = []
epoch_d_loss = []
for batch in train_loader_private:
real_data = batch[0].to(device)
batch_size = real_data.shape[0]
# === Train Discriminator with DP ===
# Generate fake data (detached - no generator gradients)
z = torch.randn(batch_size, CONFIG["latent_dim"], device=device)
with torch.no_grad():
fake_data = generator(z)
# Discriminator predictions
d_real = discriminator_private(real_data)
d_fake = discriminator_private(fake_data)
# Wasserstein-style loss (more stable than BCE for GANs)
d_loss = -torch.mean(d_real) + torch.mean(d_fake)
opt_d.zero_grad()
d_loss.backward()
opt_d.step() # Opacus handles gradient clipping + noise internally
# After the step, so the constraint holds on the updated weights.
clip_value = 0.01
with torch.no_grad():
for p in discriminator_private._module.parameters():
p.data.clamp_(-clip_value, clip_value)
epoch_d_loss.append(d_loss.item())
# Sync discriminator weights to clean copy after each batch
# This ensures generator trains against up-to-date discriminator
discriminator_for_g.load_state_dict(discriminator_private._module.state_dict())
# === Train Generator (non-private) ===
# Use the clean discriminator copy (no DP hooks)
z = torch.randn(batch_size, CONFIG["latent_dim"], device=device)
fake_data = generator(z)
# Backprop through clean discriminator
d_fake = discriminator_for_g(fake_data)
g_loss = -torch.mean(d_fake) # Generator wants D(fake) to be high
opt_g.zero_grad()
g_loss.backward()
opt_g.step()
epoch_g_loss.append(g_loss.item())
# Track privacy spent
epsilon_spent = privacy_engine.get_epsilon(target_delta)
history["epsilon"].append(epsilon_spent)
history["g_loss"].append(np.mean(epoch_g_loss))
history["d_loss"].append(np.mean(epoch_d_loss))
if (epoch + 1) % 5 == 0 or epoch == 0:
print(
f"Epoch {epoch + 1}/{epochs} | G: {history['g_loss'][-1]:.4f} | "
f"D: {history['d_loss'][-1]:.4f} | ε: {epsilon_spent:.2f}"
)
final_epsilon = privacy_engine.get_epsilon(target_delta)
print(f"\nFinal privacy guarantee: (ε={final_epsilon:.2f}, δ={target_delta})")
return history, discriminator_private
# %% [markdown]
# ### Run Training
#
# Train the DP-GAN with the configured privacy budget. Opacus automatically
# calibrates the noise multiplier to spend exactly the target epsilon over
# the specified number of epochs.
# %%
# Check for existing checkpoint
checkpoint_exists = CHECKPOINT_PATH.exists()
if checkpoint_exists and not RETRAIN:
print(f"\nLoading DP-GAN from checkpoint: {CHECKPOINT_PATH}")
checkpoint = torch.load(CHECKPOINT_PATH, map_location=device, weights_only=False)
generator.load_state_dict(checkpoint["generator"])
discriminator.load_state_dict(checkpoint["discriminator"])
history = checkpoint["history"]
# Restore normalization parameters
data_mean = checkpoint["data_mean"]
data_std = checkpoint["data_std"]
print(f"Checkpoint loaded! Final ε={history['epsilon'][-1]:.2f}")
discriminator_trained = None # Not needed for generation
else:
if RETRAIN and checkpoint_exists:
print("\nRETRAIN=True, retraining despite existing checkpoint...")
else:
print("\nNo checkpoint found, training from scratch...")
history, discriminator_trained = train_dp_gan(
generator=generator,
discriminator=discriminator,
train_loader=train_loader,
epochs=CONFIG["epochs"],
target_epsilon=CONFIG["epsilon"],
target_delta=CONFIG["delta"],
max_grad_norm=CONFIG["max_grad_norm"],
)
# Save checkpoint
CHECKPOINT_PATH.parent.mkdir(parents=True, exist_ok=True)
checkpoint = {
"generator": generator.state_dict(),
"discriminator": discriminator_trained._module.state_dict(),
"history": history,
"data_mean": data_mean,
"data_std": data_std,
"config": CONFIG,
}
torch.save(checkpoint, CHECKPOINT_PATH)
print(f"\nCheckpoint saved to: {CHECKPOINT_PATH}")
# %% [markdown]
# ## 6. Generate Synthetic Data
# %%
@torch.no_grad()
def generate_samples(generator: nn.Module, n_samples: int) -> np.ndarray:
"""Generate synthetic samples and denormalize."""
generator.eval()
samples = []
batch_size = 256
for i in range(0, n_samples, batch_size):
n = min(batch_size, n_samples - i)
z = torch.randn(n, CONFIG["latent_dim"], device=device)
batch = generator(z).cpu().numpy()
samples.append(batch)
synthetic_normalized = np.vstack(samples)
# Denormalize
synthetic = synthetic_normalized * data_std + data_mean
return synthetic
# %% [markdown]
# ### Choosing the rows to compare against
#
# `real_data` rows are consecutive trading days, so its first N rows are one stretch
# of history rather than N draws from the period the generator was trained on. The
# comparisons below read `real_eval` and the sweep reads `sweep_eval`, both drawn at
# random from all of it.
# %%
eval_rng = np.random.default_rng(SEED)
N_GENERATE = min(len(real_data), 5000)
real_eval = real_data[eval_rng.choice(len(real_data), size=N_GENERATE, replace=False)]
SWEEP_EVAL_ROWS = min(1000, len(real_data))
sweep_eval = real_data[eval_rng.choice(len(real_data), size=SWEEP_EVAL_ROWS, replace=False)]
synthetic_data = generate_samples(generator, N_GENERATE)
print(f"Generated {len(synthetic_data)} synthetic samples with DP guarantees")
# %% [markdown]
# ## 7. Fidelity: Visual Comparison with PCA and t-SNE
#
# We project both real and DP-synthetic data into 2D to assess whether the
# generator covers the same regions of the data manifold despite privacy noise.
# %%
fig = plot_fidelity_comparison(
real_eval,
synthetic_data,
title="DP-GAN: real against synthetic at the spent privacy budget",
n_samples=min(1000, N_GENERATE),
)
show_with_alt(
fig,
"Two scatter panels comparing real and synthetic rows. In both, the real points "
"form a diffuse cloud while the synthetic points lie along thin, sharply defined "
"bands: a narrow horizontal streak in the PCA projection and smooth continuous "
"arcs in the t-SNE projection, with very little of the area the real points fill.",
)
# %% [markdown]
# **Interpretation**: With differential privacy, the generator adds calibrated noise
# to protect individual training records. This typically results in less tight
# overlap between real and synthetic distributions compared to non-private GANs.
# The trade-off is intentional: stronger privacy guarantees (lower ε) mean more
# noise and less fidelity. The quantitative metrics below measure this trade-off.
# %% [markdown]
# ## 8. Evaluate Synthetic Data Quality
# %%
def evaluate_quality(real: np.ndarray, synthetic: np.ndarray) -> dict:
"""Evaluate synthetic data quality."""
results = {}
# Statistical comparison per feature
print("\n" + "=" * 70)
print("REAL vs DP-SYNTHETIC COMPARISON")
print("=" * 70)
print(
f"{'Feature':<15} {'Real Mean':>12} {'Synth Mean':>12} {'Real Std':>12} {'Synth Std':>12}"
)
print("-" * 70)
for i, name in enumerate(feature_names):
real_mean = real[:, i].mean()
synth_mean = synthetic[:, i].mean()
real_std = real[:, i].std()
synth_std = synthetic[:, i].std()
print(
f"{name:<15} {real_mean:>12.4f} {synth_mean:>12.4f} {real_std:>12.4f} {synth_std:>12.4f}"
)
# The gaps are absolute, so a small average comes from many close features, never
# from cancellation. Each mean carries its maximum, which separates the two cases.
mean_gaps = np.abs(real.mean(axis=0) - synthetic.mean(axis=0))
std_gaps = np.abs(real.std(axis=0) - synthetic.std(axis=0))
results["mean_diff"] = np.mean(mean_gaps)
results["worst_feature_mean_diff"] = np.max(mean_gaps)
results["std_diff"] = np.mean(std_gaps)
results["worst_feature_std_diff"] = np.max(std_gaps)
# Correlation preservation (handle NaN from zero-variance features)
corr_real = np.corrcoef(real, rowvar=False)
corr_synth = np.corrcoef(synthetic, rowvar=False)
corr_real = np.nan_to_num(corr_real, nan=0.0)
corr_synth = np.nan_to_num(corr_synth, nan=0.0)
results["corr_diff"] = np.linalg.norm(corr_real - corr_synth, "fro") / corr_real.size
print("-" * 70)
print(
f"Mean absolute difference: {results['mean_diff']:.4f}"
f" (worst single feature: {results['worst_feature_mean_diff']:.4f})"
)
print(
f"Std absolute difference: {results['std_diff']:.4f}"
f" (worst single feature: {results['worst_feature_std_diff']:.4f})"
)
print(f"Correlation distance (Frobenius): {results['corr_diff']:.4f}")
return results
# %% [markdown]
# ### Run Quality Evaluation
# %%
quality = evaluate_quality(real_eval, synthetic_data)
# %% [markdown]
# ## 9. Visualize Training and Results
# %%
# Training curves with privacy tracking
fig = make_subplots(
rows=1, cols=3, subplot_titles=["Generator Loss", "Discriminator Loss", "Privacy Budget (ε)"]
)
fig.add_trace(
go.Scatter(y=history["g_loss"], mode="lines", name="G Loss", line_color=COLORS["amber"]),
row=1,
col=1,
)
fig.add_trace(
go.Scatter(y=history["d_loss"], mode="lines", name="D Loss", line_color=COLORS["blue"]),
row=1,
col=2,
)
fig.add_trace(
go.Scatter(y=history["epsilon"], mode="lines", name="ε spent", line_color=COLORS["copper"]),
row=1,
col=3,
)
fig.update_xaxes(title_text="Epoch")
fig.update_layout(
title="Generator loss, discriminator loss and spent budget by epoch",
height=350,
showlegend=False,
template="ml4t",
)
show_plotly_with_alt(
fig,
"Three panels against epoch. Generator loss swings up and down across zero for the "
"whole run without settling. Discriminator loss drops from its starting value "
"within the first epoch and stays flat near zero after that. The spent privacy "
"budget climbs across the whole run, steeply at first and then more gently, "
"reaching its target at the last epoch.",
)
# %% [markdown]
# **Interpretation**: The loss curves reveal the tension between adversarial training and
# privacy noise. Unlike standard GANs where losses stabilize, DP-GAN losses remain noisier
# because Opacus injects calibrated Gaussian noise into every gradient step. The epsilon
# trajectory in the right panel shows privacy budget consumption -- it increases monotonically
# because each training step spends a fraction of the total budget. Faster epsilon growth
# means more information is being extracted from the real data per step.
# %%
# Distribution comparison
n_features_to_plot = min(6, len(feature_names))
fig = make_subplots(
rows=2, cols=3, subplot_titles=[feature_names[i] for i in range(n_features_to_plot)]
)
positions = [(1, 1), (1, 2), (1, 3), (2, 1), (2, 2), (2, 3)]
for i, (row, col) in enumerate(positions[:n_features_to_plot]):
fig.add_trace(
go.Histogram(
x=real_data[:, i],
name="Real",
opacity=0.6,
marker_color=COLORS["blue"],
nbinsx=30,
legendgroup="real",
showlegend=(i == 0),
),
row=row,
col=col,
)
fig.add_trace(
go.Histogram(
x=synthetic_data[:, i],
name="DP Synthetic",
opacity=0.6,
marker_color=COLORS["amber"],
nbinsx=30,
legendgroup="synthetic",
showlegend=(i == 0),
),
row=row,
col=col,
)
fig.update_layout(
title="Real against DP-synthetic distributions at the spent budget",
height=500,
showlegend=True,
legend=dict(orientation="h", yanchor="bottom", y=-0.15, xanchor="center", x=0.5),
barmode="overlay",
template="ml4t",
)
show_plotly_with_alt(
fig,
"Six overlaid histogram panels, one per feature. In each, the real distribution is "
"a tall concentrated peak and the DP-synthetic distribution is much lower and "
"flatter; on several features it also spreads over a wider range than the real "
"one. On volatility the synthetic mass sits away from the real peak, and on "
"momentum the two follow each other most closely.",
)
# %%
# Correlation heatmaps
fig = make_subplots(
rows=1, cols=2, subplot_titles=["Real Data Correlations", "DP-Synthetic Correlations"]
)
corr_real = np.corrcoef(real_data, rowvar=False)
corr_synth = np.corrcoef(synthetic_data, rowvar=False)
fig.add_trace(
go.Heatmap(z=corr_real, x=feature_names, y=feature_names, colorscale="RdBu", zmid=0),
row=1,
col=1,
)
fig.add_trace(
go.Heatmap(z=corr_synth, x=feature_names, y=feature_names, colorscale="RdBu", zmid=0),
row=1,
col=2,
)
fig.update_layout(
title="Feature correlation matrices, real and DP-synthetic", height=400, template="ml4t"
)
show_plotly_with_alt(
fig,
"Two correlation matrices side by side on a shared red-to-blue scale from minus "
"one to one. The real matrix is mostly pale, with weak correlations off the "
"diagonal. The DP-synthetic matrix is saturated: nearly every off-diagonal cell is "
"at the extreme red or extreme blue end of the scale.",
)
# %% [markdown]
# **Interpretation**: this figure is about the generator trained above, at the budget
# this notebook spends. Read the two matrices against the same colour scale. The real one
# is mostly pale, because these features are only weakly correlated with each other. The
# synthetic one is saturated almost everywhere off the diagonal, every pair sitting at
# one end of the scale or the other.
#
# That is not a weak version of the real structure, it is a different kind of object.
# Correlations of plus or minus one mean each feature is close to a deterministic
# function of the others, which is what a generator produces once it has collapsed onto
# a low-dimensional set instead of a distribution. The fidelity figure above shows the
# same thing geometrically: the synthetic points lie along thin bands and arcs rather
# than filling the space the real points occupy.
#
# So the failure here is not that privacy noise drowned out the subtler correlations
# while the strong ones survived. Under this budget and this much training the generator
# has stopped producing a spread at all, and every correlation it reports is an artifact
# of that collapse. Keep this in view when reading the sweep below: a summary distance
# between two correlation matrices can sit in a narrow range while one of the two
# matrices is degenerate.
# %% [markdown]
# ## 10. Privacy-Utility Trade-off Analysis
#
# Let's see how different privacy budgets affect synthetic data quality.
# %%
# Run trade-off analysis
epsilon_values = [1.0, 5.0, 10.0, 50.0]
tradeoff_results = []
def collapse_diagnostics(sample: np.ndarray, scale: np.ndarray) -> dict:
"""Two numbers that separate a spread-out sample from a collapsed one.
A generator that has collapsed onto a low-dimensional set makes each feature close
to a deterministic function of the others, which shows up as off-diagonal
correlations near one and as a handful of singular values carrying all the variance.
Neither is visible in a distance between two correlation matrices.
Args:
sample: Rows in feature units, shape (n_rows, n_features)
scale: Per-feature standard deviation of the real data, shape (n_features,)
The variance spectrum is not scale free, and these features are not on one scale:
an annualized volatility, a daily return and a dimensionless ratio differ by orders
of magnitude. On raw units the widest feature would carry almost all the variance
in any sample, collapsed or not, so both samples are divided by the same real
standard deviations first. The correlations need no such treatment.
"""
corr = np.corrcoef(sample, rowvar=False)
off_diagonal = ~np.eye(corr.shape[0], dtype=bool)
standardized = (sample - sample.mean(axis=0)) / scale
variance = np.linalg.svd(standardized, compute_uv=False) ** 2
share = np.cumsum(variance) / variance.sum()
return {
"mean_abs_offdiag_corr": float(np.abs(corr[off_diagonal]).mean()),
"dims_for_99pct_variance": int(np.searchsorted(share, 0.99) + 1),
}
# A feature the real data holds constant carries no variance to apportion, so it is
# given unit scale rather than dividing by zero.
real_feature_scale = np.where(sweep_eval.std(axis=0) > 0, sweep_eval.std(axis=0), 1.0)
print("\n" + "=" * 60)
print("PRIVACY-UTILITY TRADE-OFF ANALYSIS")
print("=" * 60)
_real_collapse = collapse_diagnostics(sweep_eval, real_feature_scale)
print(
f"Real data, for reference: mean |off-diagonal corr| "
f"{_real_collapse['mean_abs_offdiag_corr']:.3f}, "
f"dims for 99% of variance {_real_collapse['dims_for_99pct_variance']}"
)
for eps in epsilon_values:
print(f"\n--- Testing ε = {eps} ---")
# Fresh models
gen = Generator(CONFIG["latent_dim"], CONFIG["hidden_dim"], input_dim).to(device)
disc = Discriminator(input_dim, CONFIG["hidden_dim"]).to(device)
disc = ModuleValidator.fix(disc).to(device)
# Quick training (fewer epochs)
_, _ = train_dp_gan(
generator=gen,
discriminator=disc,
train_loader=train_loader,
epochs=10,
target_epsilon=eps,
target_delta=CONFIG["delta"],
max_grad_norm=CONFIG["max_grad_norm"],
)
# Evaluate
synth = generate_samples(gen, SWEEP_EVAL_ROWS)
q = evaluate_quality(sweep_eval, synth)
collapse = collapse_diagnostics(synth, real_feature_scale)
tradeoff_results.append({"epsilon": eps, **q, **collapse})
print(
f" mean |off-diagonal corr|: {collapse['mean_abs_offdiag_corr']:.3f}"
f" dims for 99% of variance: {collapse['dims_for_99pct_variance']}"
)
# %%
results_df = pl.DataFrame(tradeoff_results)
# The whole sweep in one table, so the collapse diagnostics and the largest single
# feature gap sit beside the distances the figure plots.
print(
results_df.select(
"epsilon",
"mean_diff",
"worst_feature_mean_diff",
"corr_diff",
"mean_abs_offdiag_corr",
"dims_for_99pct_variance",
)
)
fig = make_subplots(
rows=1,
cols=2,
subplot_titles=["Mean Difference vs Privacy", "Correlation Distance vs Privacy"],
)
fig.add_trace(
go.Scatter(
x=results_df["epsilon"].to_list(),
y=results_df["mean_diff"].to_list(),
mode="lines+markers",
marker=dict(size=10),
line_color=COLORS["blue"],
),
row=1,
col=1,
)
fig.add_trace(
go.Scatter(
x=results_df["epsilon"].to_list(),
y=results_df["corr_diff"].to_list(),
mode="lines+markers",
marker=dict(size=10),
line_color=COLORS["amber"],
),
row=1,
col=2,
)
fig.update_xaxes(
title_text="Privacy Budget (ε)",
type="log",
tickvals=[1, 5, 10, 50],
ticktext=["1", "5", "10", "50"],
)
fig.update_yaxes(title_text="Error (lower is better)")
fig.update_layout(
title="Mean difference and correlation distance by privacy budget",
height=400,
showlegend=False,
template="ml4t",
)
show_plotly_with_alt(
fig,
"Two line panels against the privacy budget on a logarithmic axis. Mean difference "
"starts high at the tightest budget, falls sharply to the next one, and then moves "
"up and down within a much smaller range across the looser budgets, without "
"returning near its starting height. Correlation distance moves within a narrow "
"range and is not monotonic: it rises, drops to its lowest point, then rises again.",
)
# %% [markdown]
# **Interpretation**: differential privacy trades utility for guarantees - a tighter
# budget means more noise per gradient step, and the distribution that reaches the other
# side of it is coarser. What the sweep shows is one large effect and a great deal of noise around it.
#
# The large effect is at the tightest budget, where the mean absolute difference is
# several times its value anywhere else. That much is clear in the left panel and is
# the part worth carrying away.
#
# Across the looser budgets the same measure lands in a narrow band with no clean
# ordering, and the correlation distance in the right panel is not even monotonic: it
# rises, falls to its lowest value, and rises again. Read the printed table for the
# figures. Differences that small sit inside the run-to-run variation of a single short
# training run per budget, so this sweep cannot resolve a monotonic curve, and drawing a
# smooth "quality improves with epsilon" line through it would be reading the noise.
# Separating the moderate budgets would take several independent runs at each, averaged.
#
# One caution carried forward from the correlation matrices: both measures here are
# summaries of a difference, and a difference between two correlation matrices cannot
# say whether either of them describes a real spread. That is why the sweep also prints
# `mean_abs_offdiag_corr` and `dims_for_99pct_variance` per budget, with the real data's
# values above them for reference. Both samples are put on the real data's feature
# scales first, so the dimension count is a statement about the shape of the sample and
# not about which feature happens to be measured in the largest units. Compare each
# budget's pair against the real one: an off-diagonal correlation near one, or a
# variance the real data spreads over many directions concentrated into a couple, is a
# generator that has collapsed rather than one that is merely noisy.
# Where that is what the numbers show, a flat correlation-distance curve says the metric
# does not separate these settings, not that the settings are equally good.
#
# In regulated settings - sharing client trading data, say - a ceiling on the privacy
# budget is often imposed regardless of where the utility optimum falls.
# %% [markdown]
# ## Key Takeaways
#
# 1. **Privacy-utility tradeoff is real but dominated by the strict-privacy regime**:
# A tighter budget injects more noise into the gradients and coarsens the synthetic
# data. In this sweep the cost is large and unambiguous at the tightest budget, and
# the looser ones fall into a narrow band with no clean ordering. One short run per
# budget cannot resolve a monotonic curve, so the defensible reading is that a very
# tight budget is expensive and the moderate ones are not separated here; separating
# them would take several runs at each, averaged.
#
# 2. **Per-sample gradients via Opacus are essential**: Standard PyTorch clips the
# *batch* gradient, which does not bound any individual sample's contribution.
# Opacus computes per-sample gradients using `functorch`, clips each independently
# to `max_grad_norm`, then adds calibrated noise. Manual implementations of this
# are notoriously error-prone -- Opacus makes it correct by construction.
#
# 3. **Epsilon measures cumulative information leakage**: Each training step consumes
# a fraction of the privacy budget. The privacy accountant (Renyi DP) tracks this
# precisely. Once epsilon is spent, no further training is possible without weakening
# the guarantee. This is why `make_private_with_epsilon` calibrates noise to spread
# the budget evenly across epochs.
#
# 4. **Architecture constraints matter**: BatchNorm violates DP because it computes
# statistics across the batch, leaking information about other samples. GroupNorm
# and LayerNorm are DP-safe alternatives. Use `ModuleValidator` to check compatibility
# before training.
#
# **Book**: Section 5.8 places this notebook in the full Fidelity-Utility-Privacy framework
# that situates DP-GAN against the non-private generators in earlier notebooks.
#
# **Book**: Chapter 5, Section 5.7 discusses differential privacy within the broader
# FUP framework, including membership inference attacks and nearest-neighbor privacy
# metrics that complement the formal epsilon guarantee.
# %%
# Save DP synthetic data (consistent with other generators)
output_dir = get_output_dir(5, "dp_gan")
output_dir.mkdir(parents=True, exist_ok=True)
# Save as parquet with feature names
synthetic_df = pl.DataFrame(synthetic_data, schema=feature_names)
output_path = output_dir / "dp_gan_samples.parquet"
synthetic_df.write_parquet(output_path)
print(f"\nSaved DP synthetic data to {output_path}")
# Also save metadata
metadata = {
"epsilon": float(history["epsilon"][-1]),
"delta": CONFIG["delta"],
"n_samples": len(synthetic_data),
"features": feature_names,
"quality": quality,
}
print(f"Final privacy guarantee: (ε={metadata['epsilon']:.2f}, δ={metadata['delta']})")
print("\nDP-GAN notebook complete!")
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。