跳至正文
返回文库全部文档

用双幂变差和Lee–Mykland检验检测日内跳跃

代码 《交易机器学习》

总结

本笔记将日内股票收益中的连续价格变动与离散跳跃分开。它根据对数收益平方估算日实现方差,并使用双幂变差作为对跳跃稳健的连续部分估计;两者的非负差值用于估算跳跃方差。随后,Lee–Mykland检验使用同一交易时段内此前观测值估算的局部波动率,对各收益进行标准化。其基于Gumbel分布的临界值考虑到一天内对多个分钟数据点进行检验,并能适应局部波动率变化。

分析将该方法与固定收益阈值进行比较,描述检测到的跳跃分布和发生时点,并构建每日特征,例如跳跃次数、带符号跳跃方差和跳跃方差占实现方差的比例。示例使用少量NASDAQ-100只股票的一年常规交易时段分钟线数据。分析排除了隔夜收益;交易时段早段的分钟数据点缺少足够历史数据来估算局部波动率,分析也未考虑预定新闻。检测到的跳跃特征是后续研究的输入,并非策略盈利的证据。

核心观点

  • 实现方差同时包含连续价格变动和跳跃平方;双幂变差则能稳健估计连续部分,对孤立的大幅收益具有稳健性。
  • 实现方差与双幂变差之差截断至零后,可用于估算跳跃所致方差。
  • Lee–Mykland使用同一交易时段内此前观测值估算的局部波动率标准化收益,并采用覆盖全时段的多重检验阈值。
  • 固定收益阈值可能在平静市场状态下过度检测,也可能在背景波动率较高时漏掉跳跃。
  • 所得每日跳跃特征可用于后续标签构建和波动率分析,但需考虑交易时段覆盖情况以及未根据预定新闻进行条件调整的限制。

标签

全文
# 18_algoseek_jump_detection.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Detecting Price Jumps in Intraday Returns
#
# **Chapter 3: Market Microstructure**
#
# **Docker image**: `ml4t`
#
# ## Purpose
#
# Realized intraday variance mixes two distinct processes: a continuous
# diffusive component driven by the steady arrival of small order-flow
# innovations, and a discrete jump component driven by news, scheduled
# announcements, and large block trades. This notebook separates the two
# on AlgoSeek NASDAQ-100 minute bars using bipower variation
# (Barndorff-Nielsen and Shephard, 2004) for the continuous part and the
# Lee–Mykland (2008) nonparametric test to time individual jumps. The
# resulting daily features — jump count, signed jump variance, jump share
# of realized variance — feed directly into the label-conditioning logic
# of Chapter 7 and the volatility-regime features of Chapter 8.
#
# ## Learning Objectives
#
# - Decompose daily realized variance into continuous and jump components
#   via bipower variation
# - Apply the Lee–Mykland test to flag individual jumps with a
#   leakage-free, multiple-testing-aware critical value
# - Quantify how concentrated jumps are in time of day and across the
#   trading year
# - Show why a naive |z|>k threshold over-detects in low-vol regimes and
#   under-detects in high-vol regimes
# - Materialize a per-symbol, per-day jump-feature panel ready for
#   downstream feature engineering
#
# ## Book reference
#
# Section §3.5, *Detecting Price Jumps in Intraday Returns*.
#
# ## Prerequisites
#
# - AlgoSeek TAQ minute bars under
#   `data/equities/market/nasdaq100/minute_bars/year=YYYY/month=MM.parquet`
#   (Hive-partitioned), loaded via `load_nasdaq100_bars`.
# - Companion: `13_algoseek_minute_bars_eda` for raw-bar microstructure
#   schema.

# %%
"""Lee–Mykland jump detection on AlgoSeek minute bars."""

from __future__ import annotations

import math

import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from scipy.stats import norm

from data import load_nasdaq100_bars
from utils.paths import display_path, get_output_dir
from utils.style import COLORS, show_with_alt

# %% [markdown]
# ### Declared parameters
#
# `SYMBOLS` names three NASDAQ-100 members. AlgoSeek stores the ticker each name carried
# at the time, so Meta appears as "FB" for 2020 dates - the change to "META" came in
# 2022, and asking for the later symbol on an earlier date returns nothing.
#
# `BAR_MINUTES` sets the sampling frequency. Jump tests want bars short enough that a
# jump is not averaged away with the diffusion around it, and long enough that the
# bid-ask bounce does not dominate the return.
#
# `LEE_MYKLAND_K` is the window, in bars, over which local volatility is estimated. The
# test asks whether a return is large relative to volatility *nearby*, so the window has
# to be long enough to estimate that and short enough to stay local.
#
# `JUMP_ALPHA` is the significance level for the whole session, not for one bar. A
# session holds dozens of bars, so testing each one at that level independently would
# flag several a day by chance alone; the threshold computed below is the one that
# holds the error rate at this level across the session as a whole.

# %% tags=["parameters"]
SYMBOLS = ["AMD", "AMZN", "FB"]
START_DATE = "2020-01-01"
END_DATE = "2020-12-31"
BAR_MINUTES = 5  # Bar frequency in minutes
LEE_MYKLAND_K = 12  # Within-session local-volatility window (~60 min at 5-min)
JUMP_ALPHA = 0.01  # Family-wise significance level for the daily jump test
NAIVE_Z_THRESHOLD = 4.0  # Naive comparison rule: flag |return / rolling_std| > NAIVE_Z

# %%
OUTPUT_DIR = get_output_dir(3, "jump_features")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)

# Bipower variation theoretical constant: E[|Z|] = sqrt(2/pi) for Z ~ N(0,1)
MU1 = math.sqrt(2.0 / math.pi)
BARS_PER_DAY = (16 - 9.5) * 60 // BAR_MINUTES  # Regular hours, 78 at 5-min

# %% [markdown]
# ## 1. Load minute bars and compute log returns
#
# We restrict to regular hours (09:30–16:00 ET) so that the diffusion
# assumption underlying both bipower variation and the Lee–Mykland test
# is not contaminated by extended-session stale quotes. Log returns are
# computed per symbol-day to avoid bridging the overnight gap, which
# behaves as a jump but is not the object of intraday inference.

# %%
bars = (
    load_nasdaq100_bars(
        frequency=f"{BAR_MINUTES}m",
        symbols=SYMBOLS,
        start_date=START_DATE,
        end_date=END_DATE,
        regular_hours=True,
    )
    .sort(["symbol", "timestamp"])
    .with_columns(date=pl.col("timestamp").dt.date())
)
per_symbol_days = bars.group_by("symbol").agg(n_days=pl.col("date").n_unique())
missing = [s for s in SYMBOLS if s not in set(per_symbol_days["symbol"].to_list())]
if missing:
    raise ValueError(
        f"load_nasdaq100_bars returned zero rows for {missing}; check the ticker "
        f"history (e.g. FB→META in mid-2022) against the date window."
    )
max_days = per_symbol_days["n_days"].max()
short = per_symbol_days.filter(pl.col("n_days") < 0.5 * max_days)
if short.height:
    raise ValueError(
        f"Coverage gap: {short.to_dicts()} cover <50% of the longest-symbol session "
        f"count ({max_days}); a window-shift bug may have left a ticker partially "
        f"unavailable in the requested range."
    )

returns = bars.with_columns(
    log_return=(pl.col("close").log() - pl.col("close").log().shift(1)).over(["symbol", "date"]),
).drop_nulls("log_return")

print(f"Bars loaded: {bars.height:,}  symbols: {returns['symbol'].n_unique()}")
print(f"Trading days: {returns['date'].n_unique():,}  bars/day target: {int(BARS_PER_DAY)}")

# %% [markdown]
# ## 2. Realized variance and bipower variation
#
# For day $d$ with $n$ intraday log returns $r_1, \dots, r_n$, the
# realized variance is
# $$\mathrm{RV}_d = \sum_{i=1}^{n} r_i^2,$$
# which is a consistent estimator of total quadratic variation —
# continuous *plus* squared jumps. The bipower variation,
# $$\mathrm{BV}_d = \mu_1^{-2} \sum_{i=2}^{n} |r_i| \, |r_{i-1}|, \qquad \mu_1 = \sqrt{2/\pi},$$
# is jump-robust: a single large $|r_i|$ is multiplied by its small
# neighbor and contributes negligibly. The gap
# $\max(\mathrm{RV}_d - \mathrm{BV}_d, 0)$ is therefore an estimate of
# the jump-attributable variance on day $d$.

# %%
daily = (
    returns.group_by(["symbol", "date"])
    .agg(
        rv=(pl.col("log_return") ** 2).sum(),
        bv=(pl.col("log_return").abs() * pl.col("log_return").abs().shift(1)).sum() / (MU1**2),
        n_bars=pl.len(),
    )
    .with_columns(
        jump_var=(pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0),
        jump_share_rv=((pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0) / pl.col("rv")).clip(
            0, 1
        ),
    )
)
print(f"Symbol-days: {daily.height:,}  symbols: {daily['symbol'].n_unique()}")

# %% [markdown]
# The jump share of realized variance is right-skewed across symbol-days:
# most days are almost purely diffusive (RV ≈ BV, so the jump share sits
# near zero), while a heavy right tail carries the episodic jump days. The
# ECDF makes the concentration explicit — the median day attributes little
# variance to jumps, but the upper decile can attribute a large fraction.

# %%
fig, ax = plt.subplots(figsize=(8, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
    shares = daily.filter(pl.col("symbol") == sym)["jump_share_rv"].sort().to_numpy()
    ecdf = np.arange(1, len(shares) + 1) / len(shares)
    ax.step(shares * 100, ecdf, where="post", label=sym, lw=1.6, color=color)
ax.axhline(0.5, color=COLORS["neutral"], ls="--", lw=0.8)
ax.set_xlabel("Jump share of realized variance (%)")
ax.set_ylabel("Cumulative fraction of symbol-days")
ax.set_title("Jump share of realized variance, by symbol")
ax.legend(title="Symbol")
show_with_alt(
    fig,
    "A cumulative distribution curve per symbol, stepping from zero to one, plotting the share of a symbol-day's realized variance attributed to jumps along the horizontal axis against the fraction of symbol-days at or below it, with a dashed horizontal line at the median.",
)

# %% [markdown]
# Annualized, BV gives the continuous-volatility floor of each name and
# the jump-variance share quantifies how much of the total annual
# realized variance is attributable to discrete jumps rather than
# diffusion.

# %%
annual_summary = (
    daily.group_by("symbol")
    .agg(
        rv_annual=pl.col("rv").sum(),
        bv_annual=pl.col("bv").sum(),
        jump_var_annual=pl.col("jump_var").sum(),
    )
    .with_columns(
        jump_share_annual=pl.col("jump_var_annual") / pl.col("rv_annual"),
        cont_vol_annual=(pl.col("bv_annual").sqrt()),
        rv_vol_annual=(pl.col("rv_annual").sqrt()),
    )
)
annual_summary

# %% [markdown]
# ## 3. The Lee–Mykland jump test
#
# Lee and Mykland (2008) construct a test that times jumps to the bar.
# For bar $i$ within a session of $n$ bars they form a standardized statistic
# $$L_i = \frac{r_i}{\hat\sigma_i}, \qquad
#   \hat\sigma_i^2 = \frac{1}{K-2}\sum_{j=i-K+2}^{i-1} |r_j|\,|r_{j-1}|,$$
# where $\hat\sigma_i$ is a local bipower volatility built from a window
# of $K$ prior returns *within the same session* — it adapts to slow-moving
# intraday volatility regimes and is robust to the jumps it is trying to
# detect. Building the window per session avoids bridging the overnight
# gap, which behaves as a jump and would contaminate the local-volatility
# denominator for the first $K$ bars of each day. The trade-off is that
# the first $K-1$ bars of every session are not directly testable.
#
# Under the no-jump-at-$i$ null and a continuous-diffusion alternative,
# $\max_i |L_i|$ over the daily session converges to a Gumbel distribution.
# The rejection rule at level $\alpha$ is
# $$|L_i| > S_n \beta^*(\alpha) + C_n, \qquad
#   C_n = \frac{(2\log n)^{1/2}}{\mu_1} - \frac{\log\pi + \log\log n}{2\mu_1 (2\log n)^{1/2}}, \;
#   S_n = \frac{1}{\mu_1 (2\log n)^{1/2}},$$
# and $\beta^*(\alpha) = -\log(-\log(1-\alpha))$. The multiple-testing
# burden is absorbed by the $n$-dependent constants, not by an ad-hoc
# Bonferroni adjustment. Lee and Mykland recommend $K \approx \sqrt n$
# (≈9 for 78 5-minute bars per day); we use $K = 12$ as a slightly more
# conservative window.


# %%
def gumbel_threshold(n: int, alpha: float) -> float:
    """Lee–Mykland critical value for n bars per session at level alpha."""
    c_n = (np.sqrt(2 * np.log(n)) / MU1) - (np.log(np.pi) + np.log(np.log(n))) / (
        2 * MU1 * np.sqrt(2 * np.log(n))
    )
    s_n = 1.0 / (MU1 * np.sqrt(2 * np.log(n)))
    beta_star = -np.log(-np.log(1 - alpha))
    return s_n * beta_star + c_n


def lee_mykland_jumps_session(rets: pl.DataFrame, k: int, threshold: float) -> pl.DataFrame:
    """Apply Lee–Mykland (2008) per (symbol, date) so the rolling window
    of K-2 pair products is built only from same-session bars.

    Returns the input frame with three columns appended: L_stat,
    sigma_local, is_jump. The first K-1 bars of each session have NaN
    sigma_local (no within-session history yet) and is_jump=False.
    """
    sessions: list[pl.DataFrame] = []
    for (_sym, _date), session in rets.group_by(["symbol", "date"], maintain_order=True):
        r = session["log_return"].to_numpy()
        n_obs = r.shape[0]
        sigma_hat = np.full(n_obs, np.nan)
        if n_obs >= k:
            abs_r = np.abs(r)
            # pair_prod[j] = |r_{j}| * |r_{j-1}| for j = 1 .. n_obs-1, same-session
            pair_prod = abs_r[1:] * abs_r[:-1]
            ps_cum = np.concatenate(([0.0], np.cumsum(pair_prod)))
            # For bar i (i >= k-1, 0-indexed), window uses pair_prod[i-k+1 .. i-2]
            # (k-2 terms ending at bar i-1); first testable bar is index k-1, so
            # exactly K-1 leading bars per session carry NaN sigma_local.
            idx = np.arange(k - 1, n_obs)
            sigma_sq = (ps_cum[idx - 1] - ps_cum[idx - k + 1]) / (k - 2)
            sigma_hat[idx] = np.sqrt(np.maximum(sigma_sq, 0.0))
        # NaN sigma_local → NaN L → comparison yields False; no explicit guard needed.
        L = r / sigma_hat
        sessions.append(
            session.with_columns(
                L_stat=pl.Series(L),
                sigma_local=pl.Series(sigma_hat),
                is_jump=pl.Series(np.abs(L) > threshold),
            )
        )
    return pl.concat(sessions)


# %% [markdown]
# The threshold depends on how many bars a session holds, because testing more bars gives
# more chances for a large return to appear by accident. It is computed once from the
# median session length so that every bar in the study is tested against the same
# threshold, rather than a symbol with a slightly longer session facing a slightly
# harder test.

# %%
n_per_session = returns.group_by(["symbol", "date"]).len()["len"].to_numpy()
N_PER_SESSION = int(np.median(n_per_session))
threshold_used = gumbel_threshold(N_PER_SESSION, JUMP_ALPHA)
flagged = lee_mykland_jumps_session(returns, k=LEE_MYKLAND_K, threshold=threshold_used)
jumps = flagged.filter(pl.col("is_jump"))
print(f"Bars per session (median): {N_PER_SESSION}")
print(f"Lee–Mykland critical value at α={JUMP_ALPHA}: {threshold_used:.3f}")
print(f"Jumps flagged: {jumps.height:,} across {flagged['symbol'].n_unique()} symbols")
print(
    f"Jumps per symbol-day: {jumps.height / max(1, flagged.group_by(['symbol', 'date']).len().height):.2f}"
)

# %% [markdown]
# ## 4. How often do jumps fire and how big are they?
#
# A symbol-day typically sees a handful of test rejections; days with
# scheduled news cluster materially higher. The signed magnitude of
# the underlying log return distinguishes upside surprises from
# downside.

# %%
daily_jump_counts = (
    flagged.group_by(["symbol", "date"]).agg(
        jump_count=pl.col("is_jump").sum(),
        signed_jump_var=pl.when(pl.col("is_jump"))
        .then(pl.col("log_return") ** 2 * pl.col("log_return").sign())
        .otherwise(0.0)
        .sum(),
    )
).join(
    daily.select("symbol", "date", "rv", "bv", "jump_var", "jump_share_rv"), on=["symbol", "date"]
)

# %%
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
    counts = daily_jump_counts.filter(pl.col("symbol") == sym)["jump_count"].to_numpy()
    axes[0].hist(counts, bins=range(0, int(counts.max()) + 2), alpha=0.5, label=sym, color=color)
axes[0].set_xlabel("Lee–Mykland jumps per day (count)")
axes[0].set_ylabel("Trading days")
axes[0].set_title("Lee-Mykland jumps detected per trading day")
axes[0].legend(title="Symbol")

# Local volatility is undefined for the first bars of each session, before the window
# has filled, so those returns standardize to NaN and are dropped explicitly.
all_z_raw = (flagged["log_return"] / flagged["sigma_local"]).to_numpy()
all_z = all_z_raw[np.isfinite(all_z_raw)]
filt = flagged.filter(~pl.col("is_jump"))
filt_z_raw = (filt["log_return"] / filt["sigma_local"]).to_numpy()
filt_z = filt_z_raw[np.isfinite(filt_z_raw)]
qs = np.linspace(0.001, 0.999, 200)
norm_q = norm.ppf(qs)
axes[1].scatter(
    norm_q, np.quantile(all_z, qs), s=10, alpha=0.7, label="All returns", color=COLORS["blue"]
)
axes[1].scatter(
    norm_q, np.quantile(filt_z, qs), s=10, alpha=0.7, label="Jump-filtered", color=COLORS["amber"]
)
axes[1].plot([-4, 4], [-4, 4], ls="--", lw=0.8, color=COLORS["neutral"])
axes[1].set_xlim(-4, 4)
axes[1].set_ylim(-8, 8)
axes[1].set_xlabel("Theoretical quantile, N(0,1) (std. dev.)")
axes[1].set_ylabel("Empirical quantile (std. dev.)")
axes[1].set_title("Standardized returns against the normal, with and without flagged jumps")
axes[1].legend()
show_with_alt(
    fig,
    "Two panels. The left is an overlaid histogram, one series per symbol, of how many jumps were detected per trading day. The right is a quantile-quantile plot of standardized returns against the standard normal, with one series for all returns and another for returns with flagged jumps removed, and a dashed forty-five degree reference line.",
)

# %% [markdown]
# ## 5. Jumps cluster in time of day
#
# The opening minutes and the close concentrate price discovery: news
# released overnight is impounded at the auction, and end-of-day
# rebalancing flows produce sharp moves. The session-boundary fix means
# the first `LEE_MYKLAND_K - 1` bars of each day (≈55 min) are not
# testable, so any "first 30 min" detection here is limited; we report
# the share within the testable window and look for the close-of-day
# uptick that the test can see in full.

# %%
tod = flagged.filter(pl.col("is_jump")).with_columns(
    minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
    + pl.col("timestamp").dt.minute()
)
tod_hist = tod.group_by("minute_of_day").len().sort("minute_of_day")
all_minutes = (
    flagged.with_columns(
        minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
        + pl.col("timestamp").dt.minute()
    )
    .group_by("minute_of_day")
    .len()
    .rename({"len": "all_bars"})
    .sort("minute_of_day")
)
tod_rate = tod_hist.join(all_minutes, on="minute_of_day", how="right").with_columns(
    rate_per_1k=pl.col("len").fill_null(0) / pl.col("all_bars") * 1000
)
share_first_30 = tod.filter(pl.col("minute_of_day") < 9 * 60 + 30 + 30).height / tod.height
share_last_30 = tod.filter(pl.col("minute_of_day") >= 16 * 60 - 30).height / tod.height
print(f"Share of jumps in first 30 min: {share_first_30:.1%}")
print(f"Share of jumps in last 30 min:  {share_last_30:.1%}")

# %%
fig, ax = plt.subplots(figsize=(9, 3.5))
ax.bar(
    tod_rate["minute_of_day"].to_numpy() / 60,
    tod_rate["rate_per_1k"].to_numpy(),
    width=BAR_MINUTES / 60.0,
    edgecolor="none",
    color=COLORS["blue"],
)
ax.set_xlabel("Hour of day (ET)")
ax.set_ylabel("Jumps per 1,000 bars")
ax.set_title("Jump rate by time of day")
ax.axvspan(9.5, 10.0, alpha=0.12, color=COLORS["neutral"])
ax.axvspan(15.5, 16.0, alpha=0.12, color=COLORS["neutral"])
show_with_alt(
    fig,
    "A chart of the rate at which jumps are detected against time of day across the trading session.",
)

# %% [markdown]
# ## 6. Variance decomposition over the year
#
# Stacking the continuous (bipower) and jump variance components over
# time shows that jump contribution is episodic: most days have a small
# or zero jump variance, punctuated by clusters around earnings,
# macroeconomic releases, and stress periods.

# %%
fig, ax = plt.subplots(figsize=(11, 4))
focus = daily.filter(pl.col("symbol") == SYMBOLS[0]).sort("date").to_pandas()
ax.fill_between(
    focus["date"], 0, focus["bv"], alpha=0.7, color=COLORS["neutral"], label="Continuous (BV)"
)
ax.fill_between(
    focus["date"],
    focus["bv"],
    focus["bv"] + focus["jump_var"],
    alpha=0.85,
    color=COLORS["amber"],
    label="Jump variance (RV − BV)",
)
ax.set_xlabel("Date (2020)")
ax.set_ylabel("Daily realized variance (sum of squared log returns)")
ax.set_title(f"{SYMBOLS[0]}: daily jump and diffusive variance over time")
ax.legend()
show_with_alt(
    fig,
    "Two series over time for one symbol: the daily variance attributed to jumps and the daily variance attributed to continuous diffusion, on the same axis.",
)

# %% [markdown]
# ## 7. Why the local-volatility adjustment matters
#
# A naive rule that flags any bar with $|r_i / \hat\sigma_d| > 4$, where
# $\hat\sigma_d$ is a single daily volatility, ignores time-of-day
# volatility shape and slow regime changes. Compared with Lee–Mykland
# it over-detects on calm days (the daily $\hat\sigma_d$ is small, so
# even modestly elevated bars exceed the fixed $4\hat\sigma_d$
# threshold) and under-detects on volatile days (the daily $\hat\sigma_d$
# is inflated by the very moves we are trying to flag, so large bars
# fail to clear $4\hat\sigma_d$).

# %%
daily_sigma = returns.group_by(["symbol", "date"]).agg(daily_std=pl.col("log_return").std())
naive = (
    returns.join(daily_sigma, on=["symbol", "date"])
    .with_columns(naive_z=pl.col("log_return") / pl.col("daily_std"))
    .with_columns(naive_jump=pl.col("naive_z").abs() > NAIVE_Z_THRESHOLD)
)
comparison = (
    flagged.join(naive.select("symbol", "timestamp", "naive_jump"), on=["symbol", "timestamp"])
    .group_by("symbol")
    .agg(
        lm_jumps=pl.col("is_jump").sum(),
        naive_jumps=pl.col("naive_jump").sum(),
        both=(pl.col("is_jump") & pl.col("naive_jump")).sum(),
        lm_only=(pl.col("is_jump") & ~pl.col("naive_jump")).sum(),
        naive_only=(~pl.col("is_jump") & pl.col("naive_jump")).sum(),
    )
    .sort("symbol")
)

# %% [markdown]
# Split each symbol's flags into three buckets: bars both rules agree are
# jumps, bars only Lee–Mykland flags (the naive rule *misses* these, mostly
# on volatile days when the inflated daily $\hat\sigma_d$ swallows real
# jumps), and bars only the naive rule flags (it *over-fires* on calm days
# when the small daily $\hat\sigma_d$ makes ordinary bars look extreme).
# Substantial disagreement on either side means the fixed-threshold rule is
# not a safe substitute.

# %%
syms = comparison["symbol"].to_list()
x = np.arange(len(syms))
w = 0.26
fig, ax = plt.subplots(figsize=(9, 4))
ax.bar(x - w, comparison["both"].to_numpy(), w, color=COLORS["blue"], label="Both agree")
ax.bar(x, comparison["lm_only"].to_numpy(), w, color=COLORS["amber"], label="Lee–Mykland only")
ax.bar(x + w, comparison["naive_only"].to_numpy(), w, color=COLORS["copper"], label="Naive only")
ax.set_xticks(x)
ax.set_xticklabels(syms)
ax.set_xlabel("Symbol")
ax.set_ylabel("Flagged bars (count)")
ax.set_title("Bars flagged by a fixed return threshold against Lee-Mykland")
ax.legend()
show_with_alt(
    fig,
    "A comparison of which bars two rules flag as jumps: a fixed return threshold and the Lee-Mykland test, showing where they agree and where each fires alone.",
)

# %% [markdown]
# ## 8. Materialize the jump-feature panel
#
# The output is a per-symbol, per-day frame with the features that
# downstream chapters consume directly: total jump count, signed jump
# variance, and the jump share of total realized variance. These join
# cleanly to the bar-level feature panels of Chapters 7–8.

# %%
features = daily_jump_counts.select(
    "symbol",
    "date",
    "jump_count",
    "signed_jump_var",
    pl.col("jump_var").alias("jump_variance"),
    pl.col("jump_share_rv").alias("jump_share"),
    pl.col("rv").alias("rv_total"),
    pl.col("bv").alias("continuous_variance"),
).sort(["symbol", "date"])
features.write_parquet(OUTPUT_DIR / "daily_jump_features.parquet")
print(
    f"Wrote {features.height:,} rows to {display_path(OUTPUT_DIR / 'daily_jump_features.parquet')}"
)
features.head(10)

# %% [markdown]
# ## Takeaways
#
# - Realized variance over-states diffusive risk because it absorbs
#   jumps; bipower variation gives a jump-robust continuous-variance
#   estimate and the difference times jump contribution.
# - The Lee–Mykland test fires the multiple-testing-correct number of
#   jumps per day in expectation and adapts to time-varying volatility,
#   so it is comparable across regimes in a way that fixed-threshold
#   rules are not.
# - The Lee–Mykland jump rate by time of day shows a clear close-of-day
#   uptick within the testable window (the first `LEE_MYKLAND_K - 1` bars
#   of each session are not directly testable). See §5 for the printed
#   per-window shares. The notebook does not condition on scheduled news
#   days; the daily variance decomposition in §6 shows episodic jump
#   contribution but is not aligned to an external news calendar.
# - The daily features written above plug into Chapter 7 (jump-conditional
#   labels and event-time sampling) and Chapter 8 (volatility-regime
#   features and feature stratification).

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。