تشخیص جهشهای درونروزی قیمت با تغییرات بایپاور و لی–مایکلند
خلاصه
این دفترچه توضیح میدهد چگونه تغییرات پیوسته قیمت را از جهشها در بازده سهام درونروزی جدا کنیم. واریانس تحققیافته روزانه را از بازدههای مجذور برآورد میکند و از تغییرات بایپاور بهعنوان برآوردی مقاوم به جهش از مؤلفه پیوسته استفاده میکند؛ تفاضل نامنفی این دو، واریانس جهش را برآورد میکند. سپس آزمون لی–مایکلند را بهکار میگیرد و هر بازده را با برآورد نوسان محلیای استاندارد میکند که از بازدههای پیشین در همان جلسه معاملاتی ساخته شده است.
این آزمون از مقدار بحرانی کل جلسه بر پایه آماره بیشینه استفاده میکند و چندین آزمون بازه میلهای را در نظر میگیرد. دفترچه جهشهای شناساییشده را با آستانه ثابت بازده مقایسه میکند، سهم و زمانبندی جهشها را میسنجد و ویژگیهای روزانهای مانند تعداد جهش، واریانس جهش علامتدار و سهم جهش را برای مدلسازی بعدی میسازد. نمونهها از میلههای قیمتی پنجدقیقهای NASDAQ-100 برای سهام منتخب استفاده میکنند. بازههای قیمتی ابتدایی برای برآورد محلی سابقه کافی ندارند و تحلیل جهشها را با تقویم اخبار همتراز نمیکند؛ نتیجهگیریها همچنین به بسامد بازههای قیمتی و انتخاب پنجره وابستهاند.
ایدههای کلیدی
- واریانس تحققیافته هم تغییرات پیوسته و هم جهشهای مجذور را در بر میگیرد، در حالی که تغییرات بایپاور مؤلفه پیوسته را بهطور مقاوم برآورد میکند.
- آماره لی–مایکلند هر بازده را با نوسانی مقایسه میکند که از بازدههای پیشین نزدیک در همان جلسه برآورد شده است.
- مقدار بحرانی کل جلسه، آزمون چندین میله را در نظر میگیرد و تشخیص جهش را با نوسان محلی متغیر سازگار میکند.
- آستانه ثابت بازده استانداردشده ممکن است در رژیمهای نوسان متفاوت با آزمون لی–مایکلند اختلاف داشته باشد.
- ویژگیهای جهش روزانه میتوانند برای برچسبگذاری بعدی و تحلیل رژیم نوسان بهکار روند.
برچسبها
متن کامل
# Detecting Price Jumps in Intraday Returns
# Detecting Price Jumps in Intraday Returns
**Chapter 3: Market Microstructure**
**Docker image**: `ml4t`
## Purpose
Realized intraday variance mixes two distinct processes: a continuous
diffusive component driven by the steady arrival of small order-flow
innovations, and a discrete jump component driven by news, scheduled
announcements, and large block trades. This notebook separates the two
on AlgoSeek NASDAQ-100 minute bars using bipower variation
(Barndorff-Nielsen and Shephard, 2004) for the continuous part and the
Lee–Mykland (2008) nonparametric test to time individual jumps. The
resulting daily features — jump count, signed jump variance, jump share
of realized variance — feed directly into the label-conditioning logic
of Chapter 7 and the volatility-regime features of Chapter 8.
## Learning Objectives
- Decompose daily realized variance into continuous and jump components
via bipower variation
- Apply the Lee–Mykland test to flag individual jumps with a
leakage-free, multiple-testing-aware critical value
- Quantify how concentrated jumps are in time of day and across the
trading year
- Show why a naive |z|>k threshold over-detects in low-vol regimes and
under-detects in high-vol regimes
- Materialize a per-symbol, per-day jump-feature panel ready for
downstream feature engineering
## Book reference
Section §3.5, *Detecting Price Jumps in Intraday Returns*.
## Prerequisites
- AlgoSeek TAQ minute bars under
`data/equities/market/nasdaq100/minute_bars/year=YYYY/month=MM.parquet`
(Hive-partitioned), loaded via `load_nasdaq100_bars`.
schema.
```python
"""Lee–Mykland jump detection on AlgoSeek minute bars."""
from __future__ import annotations
import math
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from scipy.stats import norm
from data import load_nasdaq100_bars
from utils.paths import display_path, get_output_dir
from utils.style import COLORS, show_with_alt
```
### Declared parameters
`SYMBOLS` names three NASDAQ-100 members. AlgoSeek stores the ticker each name carried
at the time, so Meta appears as "FB" for 2020 dates - the change to "META" came in
2022, and asking for the later symbol on an earlier date returns nothing.
`BAR_MINUTES` sets the sampling frequency. Jump tests want bars short enough that a
jump is not averaged away with the diffusion around it, and long enough that the
bid-ask bounce does not dominate the return.
`LEE_MYKLAND_K` is the window, in bars, over which local volatility is estimated. The
test asks whether a return is large relative to volatility *nearby*, so the window has
to be long enough to estimate that and short enough to stay local.
`JUMP_ALPHA` is the significance level for the whole session, not for one bar. A
session holds dozens of bars, so testing each one at that level independently would
flag several a day by chance alone; the threshold computed below is the one that
holds the error rate at this level across the session as a whole.
```python
SYMBOLS = ["AMD", "AMZN", "FB"]
START_DATE = "2020-01-01"
END_DATE = "2020-12-31"
BAR_MINUTES = 5 # Bar frequency in minutes
LEE_MYKLAND_K = 12 # Within-session local-volatility window (~60 min at 5-min)
JUMP_ALPHA = 0.01 # Family-wise significance level for the daily jump test
NAIVE_Z_THRESHOLD = 4.0 # Naive comparison rule: flag |return / rolling_std| > NAIVE_Z
```
```python
OUTPUT_DIR = get_output_dir(3, "jump_features")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
# Bipower variation theoretical constant: E[|Z|] = sqrt(2/pi) for Z ~ N(0,1)
MU1 = math.sqrt(2.0 / math.pi)
BARS_PER_DAY = (16 - 9.5) * 60 // BAR_MINUTES # Regular hours, 78 at 5-min
```
## 1. Load minute bars and compute log returns
We restrict to regular hours (09:30–16:00 ET) so that the diffusion
assumption underlying both bipower variation and the Lee–Mykland test
is not contaminated by extended-session stale quotes. Log returns are
computed per symbol-day to avoid bridging the overnight gap, which
behaves as a jump but is not the object of intraday inference.
```python
bars = (
load_nasdaq100_bars(
frequency=f"{BAR_MINUTES}m",
symbols=SYMBOLS,
start_date=START_DATE,
end_date=END_DATE,
regular_hours=True,
)
.sort(["symbol", "timestamp"])
.with_columns(date=pl.col("timestamp").dt.date())
)
per_symbol_days = bars.group_by("symbol").agg(n_days=pl.col("date").n_unique())
missing = [s for s in SYMBOLS if s not in set(per_symbol_days["symbol"].to_list())]
if missing:
raise ValueError(
f"load_nasdaq100_bars returned zero rows for {missing}; check the ticker "
f"history (e.g. FB→META in mid-2022) against the date window."
)
max_days = per_symbol_days["n_days"].max()
short = per_symbol_days.filter(pl.col("n_days") < 0.5 * max_days)
if short.height:
raise ValueError(
f"Coverage gap: {short.to_dicts()} cover <50% of the longest-symbol session "
f"count ({max_days}); a window-shift bug may have left a ticker partially "
f"unavailable in the requested range."
)
returns = bars.with_columns(
log_return=(pl.col("close").log() - pl.col("close").log().shift(1)).over(["symbol", "date"]),
).drop_nulls("log_return")
print(f"Bars loaded: {bars.height:,} symbols: {returns['symbol'].n_unique()}")
print(f"Trading days: {returns['date'].n_unique():,} bars/day target: {int(BARS_PER_DAY)}")
```
## 2. Realized variance and bipower variation
For day $d$ with $n$ intraday log returns $r_1, \dots, r_n$, the
realized variance is
$$\mathrm{RV}_d = \sum_{i=1}^{n} r_i^2,$$
which is a consistent estimator of total quadratic variation —
continuous *plus* squared jumps. The bipower variation,
$$\mathrm{BV}_d = \mu_1^{-2} \sum_{i=2}^{n} |r_i| \, |r_{i-1}|, \qquad \mu_1 = \sqrt{2/\pi},$$
is jump-robust: a single large $|r_i|$ is multiplied by its small
neighbor and contributes negligibly. The gap
$\max(\mathrm{RV}_d - \mathrm{BV}_d, 0)$ is therefore an estimate of
the jump-attributable variance on day $d$.
```python
daily = (
returns.group_by(["symbol", "date"])
.agg(
rv=(pl.col("log_return") ** 2).sum(),
bv=(pl.col("log_return").abs() * pl.col("log_return").abs().shift(1)).sum() / (MU1**2),
n_bars=pl.len(),
)
.with_columns(
jump_var=(pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0),
jump_share_rv=((pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0) / pl.col("rv")).clip(
0, 1
),
)
)
print(f"Symbol-days: {daily.height:,} symbols: {daily['symbol'].n_unique()}")
```
The jump share of realized variance is right-skewed across symbol-days:
most days are almost purely diffusive (RV ≈ BV, so the jump share sits
near zero), while a heavy right tail carries the episodic jump days. The
ECDF makes the concentration explicit — the median day attributes little
variance to jumps, but the upper decile can attribute a large fraction.
```python
fig, ax = plt.subplots(figsize=(8, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
shares = daily.filter(pl.col("symbol") == sym)["jump_share_rv"].sort().to_numpy()
ecdf = np.arange(1, len(shares) + 1) / len(shares)
ax.step(shares * 100, ecdf, where="post", label=sym, lw=1.6, color=color)
ax.axhline(0.5, color=COLORS["neutral"], ls="--", lw=0.8)
ax.set_xlabel("Jump share of realized variance (%)")
ax.set_ylabel("Cumulative fraction of symbol-days")
ax.set_title("Jump share of realized variance, by symbol")
ax.legend(title="Symbol")
show_with_alt(
fig,
"A cumulative distribution curve per symbol, stepping from zero to one, plotting the share of a symbol-day's realized variance attributed to jumps along the horizontal axis against the fraction of symbol-days at or below it, with a dashed horizontal line at the median.",
)
```
Annualized, BV gives the continuous-volatility floor of each name and
the jump-variance share quantifies how much of the total annual
realized variance is attributable to discrete jumps rather than
diffusion.
```python
annual_summary = (
daily.group_by("symbol")
.agg(
rv_annual=pl.col("rv").sum(),
bv_annual=pl.col("bv").sum(),
jump_var_annual=pl.col("jump_var").sum(),
)
.with_columns(
jump_share_annual=pl.col("jump_var_annual") / pl.col("rv_annual"),
cont_vol_annual=(pl.col("bv_annual").sqrt()),
rv_vol_annual=(pl.col("rv_annual").sqrt()),
)
)
annual_summary
```
## 3. The Lee–Mykland jump test
Lee and Mykland (2008) construct a test that times jumps to the bar.
For bar $i$ within a session of $n$ bars they form a standardized statistic
$$L_i = \frac{r_i}{\hat\sigma_i}, \qquad
\hat\sigma_i^2 = \frac{1}{K-2}\sum_{j=i-K+2}^{i-1} |r_j|\,|r_{j-1}|,$$
where $\hat\sigma_i$ is a local bipower volatility built from a window
of $K$ prior returns *within the same session* — it adapts to slow-moving
intraday volatility regimes and is robust to the jumps it is trying to
detect. Building the window per session avoids bridging the overnight
gap, which behaves as a jump and would contaminate the local-volatility
denominator for the first $K$ bars of each day. The trade-off is that
the first $K-1$ bars of every session are not directly testable.
Under the no-jump-at-$i$ null and a continuous-diffusion alternative,
$\max_i |L_i|$ over the daily session converges to a Gumbel distribution.
The rejection rule at level $\alpha$ is
$$|L_i| > S_n \beta^*(\alpha) + C_n, \qquad
C_n = \frac{(2\log n)^{1/2}}{\mu_1} - \frac{\log\pi + \log\log n}{2\mu_1 (2\log n)^{1/2}}, \;
S_n = \frac{1}{\mu_1 (2\log n)^{1/2}},$$
and $\beta^*(\alpha) = -\log(-\log(1-\alpha))$. The multiple-testing
burden is absorbed by the $n$-dependent constants, not by an ad-hoc
Bonferroni adjustment. Lee and Mykland recommend $K \approx \sqrt n$
(≈9 for 78 5-minute bars per day); we use $K = 12$ as a slightly more
conservative window.
```python
def gumbel_threshold(n: int, alpha: float) -> float:
"""Lee–Mykland critical value for n bars per session at level alpha."""
c_n = (np.sqrt(2 * np.log(n)) / MU1) - (np.log(np.pi) + np.log(np.log(n))) / (
2 * MU1 * np.sqrt(2 * np.log(n))
)
s_n = 1.0 / (MU1 * np.sqrt(2 * np.log(n)))
beta_star = -np.log(-np.log(1 - alpha))
return s_n * beta_star + c_n
def lee_mykland_jumps_session(rets: pl.DataFrame, k: int, threshold: float) -> pl.DataFrame:
"""Apply Lee–Mykland (2008) per (symbol, date) so the rolling window
of K-2 pair products is built only from same-session bars.
Returns the input frame with three columns appended: L_stat,
sigma_local, is_jump. The first K-1 bars of each session have NaN
sigma_local (no within-session history yet) and is_jump=False.
"""
sessions: list[pl.DataFrame] = []
for (_sym, _date), session in rets.group_by(["symbol", "date"], maintain_order=True):
r = session["log_return"].to_numpy()
n_obs = r.shape[0]
sigma_hat = np.full(n_obs, np.nan)
if n_obs >= k:
abs_r = np.abs(r)
# pair_prod[j] = |r_{j}| * |r_{j-1}| for j = 1 .. n_obs-1, same-session
pair_prod = abs_r[1:] * abs_r[:-1]
ps_cum = np.concatenate(([0.0], np.cumsum(pair_prod)))
# For bar i (i >= k-1, 0-indexed), window uses pair_prod[i-k+1 .. i-2]
# (k-2 terms ending at bar i-1); first testable bar is index k-1, so
# exactly K-1 leading bars per session carry NaN sigma_local.
idx = np.arange(k - 1, n_obs)
sigma_sq = (ps_cum[idx - 1] - ps_cum[idx - k + 1]) / (k - 2)
sigma_hat[idx] = np.sqrt(np.maximum(sigma_sq, 0.0))
# NaN sigma_local → NaN L → comparison yields False; no explicit guard needed.
L = r / sigma_hat
sessions.append(
session.with_columns(
L_stat=pl.Series(L),
sigma_local=pl.Series(sigma_hat),
is_jump=pl.Series(np.abs(L) > threshold),
)
)
return pl.concat(sessions)
```
The threshold depends on how many bars a session holds, because testing more bars gives
more chances for a large return to appear by accident. It is computed once from the
median session length so that every bar in the study is tested against the same
threshold, rather than a symbol with a slightly longer session facing a slightly
harder test.
```python
n_per_session = returns.group_by(["symbol", "date"]).len()["len"].to_numpy()
N_PER_SESSION = int(np.median(n_per_session))
threshold_used = gumbel_threshold(N_PER_SESSION, JUMP_ALPHA)
flagged = lee_mykland_jumps_session(returns, k=LEE_MYKLAND_K, threshold=threshold_used)
jumps = flagged.filter(pl.col("is_jump"))
print(f"Bars per session (median): {N_PER_SESSION}")
print(f"Lee–Mykland critical value at α={JUMP_ALPHA}: {threshold_used:.3f}")
print(f"Jumps flagged: {jumps.height:,} across {flagged['symbol'].n_unique()} symbols")
print(
f"Jumps per symbol-day: {jumps.height / max(1, flagged.group_by(['symbol', 'date']).len().height):.2f}"
)
```
## 4. How often do jumps fire and how big are they?
A symbol-day typically sees a handful of test rejections; days with
scheduled news cluster materially higher. The signed magnitude of
the underlying log return distinguishes upside surprises from
downside.
```python
daily_jump_counts = (
flagged.group_by(["symbol", "date"]).agg(
jump_count=pl.col("is_jump").sum(),
signed_jump_var=pl.when(pl.col("is_jump"))
.then(pl.col("log_return") ** 2 * pl.col("log_return").sign())
.otherwise(0.0)
.sum(),
)
).join(
daily.select("symbol", "date", "rv", "bv", "jump_var", "jump_share_rv"), on=["symbol", "date"]
)
```
```python
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
counts = daily_jump_counts.filter(pl.col("symbol") == sym)["jump_count"].to_numpy()
axes[0].hist(counts, bins=range(0, int(counts.max()) + 2), alpha=0.5, label=sym, color=color)
axes[0].set_xlabel("Lee–Mykland jumps per day (count)")
axes[0].set_ylabel("Trading days")
axes[0].set_title("Lee-Mykland jumps detected per trading day")
axes[0].legend(title="Symbol")
# Local volatility is undefined for the first bars of each session, before the window
# has filled, so those returns standardize to NaN and are dropped explicitly.
all_z_raw = (flagged["log_return"] / flagged["sigma_local"]).to_numpy()
all_z = all_z_raw[np.isfinite(all_z_raw)]
filt = flagged.filter(~pl.col("is_jump"))
filt_z_raw = (filt["log_return"] / filt["sigma_local"]).to_numpy()
filt_z = filt_z_raw[np.isfinite(filt_z_raw)]
qs = np.linspace(0.001, 0.999, 200)
norm_q = norm.ppf(qs)
axes[1].scatter(
norm_q, np.quantile(all_z, qs), s=10, alpha=0.7, label="All returns", color=COLORS["blue"]
)
axes[1].scatter(
norm_q, np.quantile(filt_z, qs), s=10, alpha=0.7, label="Jump-filtered", color=COLORS["amber"]
)
axes[1].plot([-4, 4], [-4, 4], ls="--", lw=0.8, color=COLORS["neutral"])
axes[1].set_xlim(-4, 4)
axes[1].set_ylim(-8, 8)
axes[1].set_xlabel("Theoretical quantile, N(0,1) (std. dev.)")
axes[1].set_ylabel("Empirical quantile (std. dev.)")
axes[1].set_title("Standardized returns against the normal, with and without flagged jumps")
axes[1].legend()
show_with_alt(
fig,
"Two panels. The left is an overlaid histogram, one series per symbol, of how many jumps were detected per trading day. The right is a quantile-quantile plot of standardized returns against the standard normal, with one series for all returns and another for returns with flagged jumps removed, and a dashed forty-five degree reference line.",
)
```
## 5. Jumps cluster in time of day
The opening minutes and the close concentrate price discovery: news
released overnight is impounded at the auction, and end-of-day
rebalancing flows produce sharp moves. The session-boundary fix means
the first `LEE_MYKLAND_K - 1` bars of each day (≈55 min) are not
testable, so any "first 30 min" detection here is limited; we report
the share within the testable window and look for the close-of-day
uptick that the test can see in full.
```python
tod = flagged.filter(pl.col("is_jump")).with_columns(
minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
+ pl.col("timestamp").dt.minute()
)
tod_hist = tod.group_by("minute_of_day").len().sort("minute_of_day")
all_minutes = (
flagged.with_columns(
minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
+ pl.col("timestamp").dt.minute()
)
.group_by("minute_of_day")
.len()
.rename({"len": "all_bars"})
.sort("minute_of_day")
)
tod_rate = tod_hist.join(all_minutes, on="minute_of_day", how="right").with_columns(
rate_per_1k=pl.col("len").fill_null(0) / pl.col("all_bars") * 1000
)
share_first_30 = tod.filter(pl.col("minute_of_day") < 9 * 60 + 30 + 30).height / tod.height
share_last_30 = tod.filter(pl.col("minute_of_day") >= 16 * 60 - 30).height / tod.height
print(f"Share of jumps in first 30 min: {share_first_30:.1%}")
print(f"Share of jumps in last 30 min: {share_last_30:.1%}")
```
```python
fig, ax = plt.subplots(figsize=(9, 3.5))
ax.bar(
tod_rate["minute_of_day"].to_numpy() / 60,
tod_rate["rate_per_1k"].to_numpy(),
width=BAR_MINUTES / 60.0,
edgecolor="none",
color=COLORS["blue"],
)
ax.set_xlabel("Hour of day (ET)")
ax.set_ylabel("Jumps per 1,000 bars")
ax.set_title("Jump rate by time of day")
ax.axvspan(9.5, 10.0, alpha=0.12, color=COLORS["neutral"])
ax.axvspan(15.5, 16.0, alpha=0.12, color=COLORS["neutral"])
show_with_alt(
fig,
"A chart of the rate at which jumps are detected against time of day across the trading session.",
)
```
## 6. Variance decomposition over the year
Stacking the continuous (bipower) and jump variance components over
time shows that jump contribution is episodic: most days have a small
or zero jump variance, punctuated by clusters around earnings,
macroeconomic releases, and stress periods.
```python
fig, ax = plt.subplots(figsize=(11, 4))
focus = daily.filter(pl.col("symbol") == SYMBOLS[0]).sort("date").to_pandas()
ax.fill_between(
focus["date"], 0, focus["bv"], alpha=0.7, color=COLORS["neutral"], label="Continuous (BV)"
)
ax.fill_between(
focus["date"],
focus["bv"],
focus["bv"] + focus["jump_var"],
alpha=0.85,
color=COLORS["amber"],
label="Jump variance (RV − BV)",
)
ax.set_xlabel("Date (2020)")
ax.set_ylabel("Daily realized variance (sum of squared log returns)")
ax.set_title(f"{SYMBOLS[0]}: daily jump and diffusive variance over time")
ax.legend()
show_with_alt(
fig,
"Two series over time for one symbol: the daily variance attributed to jumps and the daily variance attributed to continuous diffusion, on the same axis.",
)
```
## 7. Why the local-volatility adjustment matters
A naive rule that flags any bar with $|r_i / \hat\sigma_d| > 4$, where
$\hat\sigma_d$ is a single daily volatility, ignores time-of-day
volatility shape and slow regime changes. Compared with Lee–Mykland
it over-detects on calm days (the daily $\hat\sigma_d$ is small, so
even modestly elevated bars exceed the fixed $4\hat\sigma_d$
threshold) and under-detects on volatile days (the daily $\hat\sigma_d$
is inflated by the very moves we are trying to flag, so large bars
fail to clear $4\hat\sigma_d$).
```python
daily_sigma = returns.group_by(["symbol", "date"]).agg(daily_std=pl.col("log_return").std())
naive = (
returns.join(daily_sigma, on=["symbol", "date"])
.with_columns(naive_z=pl.col("log_return") / pl.col("daily_std"))
.with_columns(naive_jump=pl.col("naive_z").abs() > NAIVE_Z_THRESHOLD)
)
comparison = (
flagged.join(naive.select("symbol", "timestamp", "naive_jump"), on=["symbol", "timestamp"])
.group_by("symbol")
.agg(
lm_jumps=pl.col("is_jump").sum(),
naive_jumps=pl.col("naive_jump").sum(),
both=(pl.col("is_jump") & pl.col("naive_jump")).sum(),
lm_only=(pl.col("is_jump") & ~pl.col("naive_jump")).sum(),
naive_only=(~pl.col("is_jump") & pl.col("naive_jump")).sum(),
)
.sort("symbol")
)
```
Split each symbol's flags into three buckets: bars both rules agree are
jumps, bars only Lee–Mykland flags (the naive rule *misses* these, mostly
on volatile days when the inflated daily $\hat\sigma_d$ swallows real
jumps), and bars only the naive rule flags (it *over-fires* on calm days
when the small daily $\hat\sigma_d$ makes ordinary bars look extreme).
Substantial disagreement on either side means the fixed-threshold rule is
not a safe substitute.
```python
syms = comparison["symbol"].to_list()
x = np.arange(len(syms))
w = 0.26
fig, ax = plt.subplots(figsize=(9, 4))
ax.bar(x - w, comparison["both"].to_numpy(), w, color=COLORS["blue"], label="Both agree")
ax.bar(x, comparison["lm_only"].to_numpy(), w, color=COLORS["amber"], label="Lee–Mykland only")
ax.bar(x + w, comparison["naive_only"].to_numpy(), w, color=COLORS["copper"], label="Naive only")
ax.set_xticks(x)
ax.set_xticklabels(syms)
ax.set_xlabel("Symbol")
ax.set_ylabel("Flagged bars (count)")
ax.set_title("Bars flagged by a fixed return threshold against Lee-Mykland")
ax.legend()
show_with_alt(
fig,
"A comparison of which bars two rules flag as jumps: a fixed return threshold and the Lee-Mykland test, showing where they agree and where each fires alone.",
)
```
## 8. Materialize the jump-feature panel
The output is a per-symbol, per-day frame with the features that
downstream chapters consume directly: total jump count, signed jump
variance, and the jump share of total realized variance. These join
cleanly to the bar-level feature panels of Chapters 7–8.
```python
features = daily_jump_counts.select(
"symbol",
"date",
"jump_count",
"signed_jump_var",
pl.col("jump_var").alias("jump_variance"),
pl.col("jump_share_rv").alias("jump_share"),
pl.col("rv").alias("rv_total"),
pl.col("bv").alias("continuous_variance"),
).sort(["symbol", "date"])
features.write_parquet(OUTPUT_DIR / "daily_jump_features.parquet")
print(
f"Wrote {features.height:,} rows to {display_path(OUTPUT_DIR / 'daily_jump_features.parquet')}"
)
features.head(10)
```
## Takeaways
- Realized variance over-states diffusive risk because it absorbs
jumps; bipower variation gives a jump-robust continuous-variance
estimate and the difference times jump contribution.
- The Lee–Mykland test fires the multiple-testing-correct number of
jumps per day in expectation and adapts to time-varying volatility,
so it is comparable across regimes in a way that fixed-threshold
rules are not.
- The Lee–Mykland jump rate by time of day shows a clear close-of-day
uptick within the testable window (the first `LEE_MYKLAND_K - 1` bars
of each session are not directly testable). See §5 for the printed
per-window shares. The notebook does not condition on scheduled news
days; the daily variance decomposition in §6 shows episodic jump
contribution but is not aligned to an external news calendar.
- The daily features written above plug into Chapter 7 (jump-conditional
labels and event-time sampling) and Chapter 8 (volatility-regime
features and feature stratification).




با ذکر منبع و مطابق مجوز اثر، بهطور کامل نمایش داده میشود. مجوز: MIT
این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخهای از اثر منبع نیست.