바이파워 변동과 리–마이클런드 검정으로 장중 가격 점프 탐지
노트북 Machine Learning for Trading
요약
이 노트북은 장중 주식 수익률에서 연속적인 가격 변동과 점프를 구분하는 방법을 설명합니다. 제곱 수익률로 일별 실현분산을 추정하고, 점프에 강건한 연속 성분 추정치로 바이파워 변동을 사용합니다. 두 값의 음수가 아닌 차이로 점프 분산을 추정합니다. 이어서 같은 거래 세션에서 앞서 발생한 수익률로 계산한 국소 변동성 추정치로 각 수익률을 표준화해 리–마이클런드 검정을 적용합니다.
이 검정은 최대 통계량을 바탕으로 세션 전체에 적용되는 임계값을 사용해 여러 바 검정에 따른 문제를 고려합니다. 노트북은 탐지된 점프를 고정 수익률 임계값과 비교하고, 점프 비중과 발생 시점을 살펴보며, 이후 모델링을 위한 점프 횟수, 부호 있는 점프 분산, 점프 비중 같은 일별 특성을 구성합니다. 예시에서는 선택된 주식의 5분봉 NASDAQ-100 데이터를 사용합니다. 초반 바는 국소 추정에 필요한 이력이 충분하지 않으며, 분석은 점프를 뉴스 일정과 대조하지 않습니다. 결론은 바 빈도와 윈도 설정에도 좌우됩니다.
핵심 아이디어
- 실현분산에는 연속 변동과 제곱 점프가 모두 포함되며, 바이파워 변동은 연속 성분을 강건하게 추정합니다.
- 리–마이클런드 통계량은 각 수익률을 같은 세션의 인접한 과거 수익률로 추정한 변동성과 비교합니다.
- 세션 전체에 적용되는 임계값은 여러 바를 검정하는 문제를 고려하고 변하는 국소 변동성에 맞춰 점프 탐지를 조정합니다.
- 고정된 표준화 수익률 기준값은 변동성 국면에 따라 리–마이클런드 검정과 다른 결과를 낼 수 있습니다.
- 일별 점프 특성은 후속 라벨링과 변동성 국면 분석에 활용할 수 있습니다.
태그
전문
# Detecting Price Jumps in Intraday Returns
# Detecting Price Jumps in Intraday Returns
**Chapter 3: Market Microstructure**
**Docker image**: `ml4t`
## Purpose
Realized intraday variance mixes two distinct processes: a continuous
diffusive component driven by the steady arrival of small order-flow
innovations, and a discrete jump component driven by news, scheduled
announcements, and large block trades. This notebook separates the two
on AlgoSeek NASDAQ-100 minute bars using bipower variation
(Barndorff-Nielsen and Shephard, 2004) for the continuous part and the
Lee–Mykland (2008) nonparametric test to time individual jumps. The
resulting daily features — jump count, signed jump variance, jump share
of realized variance — feed directly into the label-conditioning logic
of Chapter 7 and the volatility-regime features of Chapter 8.
## Learning Objectives
- Decompose daily realized variance into continuous and jump components
via bipower variation
- Apply the Lee–Mykland test to flag individual jumps with a
leakage-free, multiple-testing-aware critical value
- Quantify how concentrated jumps are in time of day and across the
trading year
- Show why a naive |z|>k threshold over-detects in low-vol regimes and
under-detects in high-vol regimes
- Materialize a per-symbol, per-day jump-feature panel ready for
downstream feature engineering
## Book reference
Section §3.5, *Detecting Price Jumps in Intraday Returns*.
## Prerequisites
- AlgoSeek TAQ minute bars under
`data/equities/market/nasdaq100/minute_bars/year=YYYY/month=MM.parquet`
(Hive-partitioned), loaded via `load_nasdaq100_bars`.
schema.
```python
"""Lee–Mykland jump detection on AlgoSeek minute bars."""
from __future__ import annotations
import math
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from scipy.stats import norm
from data import load_nasdaq100_bars
from utils.paths import display_path, get_output_dir
from utils.style import COLORS, show_with_alt
```
### Declared parameters
`SYMBOLS` names three NASDAQ-100 members. AlgoSeek stores the ticker each name carried
at the time, so Meta appears as "FB" for 2020 dates - the change to "META" came in
2022, and asking for the later symbol on an earlier date returns nothing.
`BAR_MINUTES` sets the sampling frequency. Jump tests want bars short enough that a
jump is not averaged away with the diffusion around it, and long enough that the
bid-ask bounce does not dominate the return.
`LEE_MYKLAND_K` is the window, in bars, over which local volatility is estimated. The
test asks whether a return is large relative to volatility *nearby*, so the window has
to be long enough to estimate that and short enough to stay local.
`JUMP_ALPHA` is the significance level for the whole session, not for one bar. A
session holds dozens of bars, so testing each one at that level independently would
flag several a day by chance alone; the threshold computed below is the one that
holds the error rate at this level across the session as a whole.
```python
SYMBOLS = ["AMD", "AMZN", "FB"]
START_DATE = "2020-01-01"
END_DATE = "2020-12-31"
BAR_MINUTES = 5 # Bar frequency in minutes
LEE_MYKLAND_K = 12 # Within-session local-volatility window (~60 min at 5-min)
JUMP_ALPHA = 0.01 # Family-wise significance level for the daily jump test
NAIVE_Z_THRESHOLD = 4.0 # Naive comparison rule: flag |return / rolling_std| > NAIVE_Z
```
```python
OUTPUT_DIR = get_output_dir(3, "jump_features")
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
# Bipower variation theoretical constant: E[|Z|] = sqrt(2/pi) for Z ~ N(0,1)
MU1 = math.sqrt(2.0 / math.pi)
BARS_PER_DAY = (16 - 9.5) * 60 // BAR_MINUTES # Regular hours, 78 at 5-min
```
## 1. Load minute bars and compute log returns
We restrict to regular hours (09:30–16:00 ET) so that the diffusion
assumption underlying both bipower variation and the Lee–Mykland test
is not contaminated by extended-session stale quotes. Log returns are
computed per symbol-day to avoid bridging the overnight gap, which
behaves as a jump but is not the object of intraday inference.
```python
bars = (
load_nasdaq100_bars(
frequency=f"{BAR_MINUTES}m",
symbols=SYMBOLS,
start_date=START_DATE,
end_date=END_DATE,
regular_hours=True,
)
.sort(["symbol", "timestamp"])
.with_columns(date=pl.col("timestamp").dt.date())
)
per_symbol_days = bars.group_by("symbol").agg(n_days=pl.col("date").n_unique())
missing = [s for s in SYMBOLS if s not in set(per_symbol_days["symbol"].to_list())]
if missing:
raise ValueError(
f"load_nasdaq100_bars returned zero rows for {missing}; check the ticker "
f"history (e.g. FB→META in mid-2022) against the date window."
)
max_days = per_symbol_days["n_days"].max()
short = per_symbol_days.filter(pl.col("n_days") < 0.5 * max_days)
if short.height:
raise ValueError(
f"Coverage gap: {short.to_dicts()} cover <50% of the longest-symbol session "
f"count ({max_days}); a window-shift bug may have left a ticker partially "
f"unavailable in the requested range."
)
returns = bars.with_columns(
log_return=(pl.col("close").log() - pl.col("close").log().shift(1)).over(["symbol", "date"]),
).drop_nulls("log_return")
print(f"Bars loaded: {bars.height:,} symbols: {returns['symbol'].n_unique()}")
print(f"Trading days: {returns['date'].n_unique():,} bars/day target: {int(BARS_PER_DAY)}")
```
## 2. Realized variance and bipower variation
For day $d$ with $n$ intraday log returns $r_1, \dots, r_n$, the
realized variance is
$$\mathrm{RV}_d = \sum_{i=1}^{n} r_i^2,$$
which is a consistent estimator of total quadratic variation —
continuous *plus* squared jumps. The bipower variation,
$$\mathrm{BV}_d = \mu_1^{-2} \sum_{i=2}^{n} |r_i| \, |r_{i-1}|, \qquad \mu_1 = \sqrt{2/\pi},$$
is jump-robust: a single large $|r_i|$ is multiplied by its small
neighbor and contributes negligibly. The gap
$\max(\mathrm{RV}_d - \mathrm{BV}_d, 0)$ is therefore an estimate of
the jump-attributable variance on day $d$.
```python
daily = (
returns.group_by(["symbol", "date"])
.agg(
rv=(pl.col("log_return") ** 2).sum(),
bv=(pl.col("log_return").abs() * pl.col("log_return").abs().shift(1)).sum() / (MU1**2),
n_bars=pl.len(),
)
.with_columns(
jump_var=(pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0),
jump_share_rv=((pl.col("rv") - pl.col("bv")).clip(lower_bound=0.0) / pl.col("rv")).clip(
0, 1
),
)
)
print(f"Symbol-days: {daily.height:,} symbols: {daily['symbol'].n_unique()}")
```
The jump share of realized variance is right-skewed across symbol-days:
most days are almost purely diffusive (RV ≈ BV, so the jump share sits
near zero), while a heavy right tail carries the episodic jump days. The
ECDF makes the concentration explicit — the median day attributes little
variance to jumps, but the upper decile can attribute a large fraction.
```python
fig, ax = plt.subplots(figsize=(8, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
shares = daily.filter(pl.col("symbol") == sym)["jump_share_rv"].sort().to_numpy()
ecdf = np.arange(1, len(shares) + 1) / len(shares)
ax.step(shares * 100, ecdf, where="post", label=sym, lw=1.6, color=color)
ax.axhline(0.5, color=COLORS["neutral"], ls="--", lw=0.8)
ax.set_xlabel("Jump share of realized variance (%)")
ax.set_ylabel("Cumulative fraction of symbol-days")
ax.set_title("Jump share of realized variance, by symbol")
ax.legend(title="Symbol")
show_with_alt(
fig,
"A cumulative distribution curve per symbol, stepping from zero to one, plotting the share of a symbol-day's realized variance attributed to jumps along the horizontal axis against the fraction of symbol-days at or below it, with a dashed horizontal line at the median.",
)
```
Annualized, BV gives the continuous-volatility floor of each name and
the jump-variance share quantifies how much of the total annual
realized variance is attributable to discrete jumps rather than
diffusion.
```python
annual_summary = (
daily.group_by("symbol")
.agg(
rv_annual=pl.col("rv").sum(),
bv_annual=pl.col("bv").sum(),
jump_var_annual=pl.col("jump_var").sum(),
)
.with_columns(
jump_share_annual=pl.col("jump_var_annual") / pl.col("rv_annual"),
cont_vol_annual=(pl.col("bv_annual").sqrt()),
rv_vol_annual=(pl.col("rv_annual").sqrt()),
)
)
annual_summary
```
## 3. The Lee–Mykland jump test
Lee and Mykland (2008) construct a test that times jumps to the bar.
For bar $i$ within a session of $n$ bars they form a standardized statistic
$$L_i = \frac{r_i}{\hat\sigma_i}, \qquad
\hat\sigma_i^2 = \frac{1}{K-2}\sum_{j=i-K+2}^{i-1} |r_j|\,|r_{j-1}|,$$
where $\hat\sigma_i$ is a local bipower volatility built from a window
of $K$ prior returns *within the same session* — it adapts to slow-moving
intraday volatility regimes and is robust to the jumps it is trying to
detect. Building the window per session avoids bridging the overnight
gap, which behaves as a jump and would contaminate the local-volatility
denominator for the first $K$ bars of each day. The trade-off is that
the first $K-1$ bars of every session are not directly testable.
Under the no-jump-at-$i$ null and a continuous-diffusion alternative,
$\max_i |L_i|$ over the daily session converges to a Gumbel distribution.
The rejection rule at level $\alpha$ is
$$|L_i| > S_n \beta^*(\alpha) + C_n, \qquad
C_n = \frac{(2\log n)^{1/2}}{\mu_1} - \frac{\log\pi + \log\log n}{2\mu_1 (2\log n)^{1/2}}, \;
S_n = \frac{1}{\mu_1 (2\log n)^{1/2}},$$
and $\beta^*(\alpha) = -\log(-\log(1-\alpha))$. The multiple-testing
burden is absorbed by the $n$-dependent constants, not by an ad-hoc
Bonferroni adjustment. Lee and Mykland recommend $K \approx \sqrt n$
(≈9 for 78 5-minute bars per day); we use $K = 12$ as a slightly more
conservative window.
```python
def gumbel_threshold(n: int, alpha: float) -> float:
"""Lee–Mykland critical value for n bars per session at level alpha."""
c_n = (np.sqrt(2 * np.log(n)) / MU1) - (np.log(np.pi) + np.log(np.log(n))) / (
2 * MU1 * np.sqrt(2 * np.log(n))
)
s_n = 1.0 / (MU1 * np.sqrt(2 * np.log(n)))
beta_star = -np.log(-np.log(1 - alpha))
return s_n * beta_star + c_n
def lee_mykland_jumps_session(rets: pl.DataFrame, k: int, threshold: float) -> pl.DataFrame:
"""Apply Lee–Mykland (2008) per (symbol, date) so the rolling window
of K-2 pair products is built only from same-session bars.
Returns the input frame with three columns appended: L_stat,
sigma_local, is_jump. The first K-1 bars of each session have NaN
sigma_local (no within-session history yet) and is_jump=False.
"""
sessions: list[pl.DataFrame] = []
for (_sym, _date), session in rets.group_by(["symbol", "date"], maintain_order=True):
r = session["log_return"].to_numpy()
n_obs = r.shape[0]
sigma_hat = np.full(n_obs, np.nan)
if n_obs >= k:
abs_r = np.abs(r)
# pair_prod[j] = |r_{j}| * |r_{j-1}| for j = 1 .. n_obs-1, same-session
pair_prod = abs_r[1:] * abs_r[:-1]
ps_cum = np.concatenate(([0.0], np.cumsum(pair_prod)))
# For bar i (i >= k-1, 0-indexed), window uses pair_prod[i-k+1 .. i-2]
# (k-2 terms ending at bar i-1); first testable bar is index k-1, so
# exactly K-1 leading bars per session carry NaN sigma_local.
idx = np.arange(k - 1, n_obs)
sigma_sq = (ps_cum[idx - 1] - ps_cum[idx - k + 1]) / (k - 2)
sigma_hat[idx] = np.sqrt(np.maximum(sigma_sq, 0.0))
# NaN sigma_local → NaN L → comparison yields False; no explicit guard needed.
L = r / sigma_hat
sessions.append(
session.with_columns(
L_stat=pl.Series(L),
sigma_local=pl.Series(sigma_hat),
is_jump=pl.Series(np.abs(L) > threshold),
)
)
return pl.concat(sessions)
```
The threshold depends on how many bars a session holds, because testing more bars gives
more chances for a large return to appear by accident. It is computed once from the
median session length so that every bar in the study is tested against the same
threshold, rather than a symbol with a slightly longer session facing a slightly
harder test.
```python
n_per_session = returns.group_by(["symbol", "date"]).len()["len"].to_numpy()
N_PER_SESSION = int(np.median(n_per_session))
threshold_used = gumbel_threshold(N_PER_SESSION, JUMP_ALPHA)
flagged = lee_mykland_jumps_session(returns, k=LEE_MYKLAND_K, threshold=threshold_used)
jumps = flagged.filter(pl.col("is_jump"))
print(f"Bars per session (median): {N_PER_SESSION}")
print(f"Lee–Mykland critical value at α={JUMP_ALPHA}: {threshold_used:.3f}")
print(f"Jumps flagged: {jumps.height:,} across {flagged['symbol'].n_unique()} symbols")
print(
f"Jumps per symbol-day: {jumps.height / max(1, flagged.group_by(['symbol', 'date']).len().height):.2f}"
)
```
## 4. How often do jumps fire and how big are they?
A symbol-day typically sees a handful of test rejections; days with
scheduled news cluster materially higher. The signed magnitude of
the underlying log return distinguishes upside surprises from
downside.
```python
daily_jump_counts = (
flagged.group_by(["symbol", "date"]).agg(
jump_count=pl.col("is_jump").sum(),
signed_jump_var=pl.when(pl.col("is_jump"))
.then(pl.col("log_return") ** 2 * pl.col("log_return").sign())
.otherwise(0.0)
.sum(),
)
).join(
daily.select("symbol", "date", "rv", "bv", "jump_var", "jump_share_rv"), on=["symbol", "date"]
)
```
```python
fig, axes = plt.subplots(1, 2, figsize=(12, 4))
sym_colors = [COLORS["blue"], COLORS["amber"], COLORS["copper"]]
for sym, color in zip(SYMBOLS, sym_colors):
counts = daily_jump_counts.filter(pl.col("symbol") == sym)["jump_count"].to_numpy()
axes[0].hist(counts, bins=range(0, int(counts.max()) + 2), alpha=0.5, label=sym, color=color)
axes[0].set_xlabel("Lee–Mykland jumps per day (count)")
axes[0].set_ylabel("Trading days")
axes[0].set_title("Lee-Mykland jumps detected per trading day")
axes[0].legend(title="Symbol")
# Local volatility is undefined for the first bars of each session, before the window
# has filled, so those returns standardize to NaN and are dropped explicitly.
all_z_raw = (flagged["log_return"] / flagged["sigma_local"]).to_numpy()
all_z = all_z_raw[np.isfinite(all_z_raw)]
filt = flagged.filter(~pl.col("is_jump"))
filt_z_raw = (filt["log_return"] / filt["sigma_local"]).to_numpy()
filt_z = filt_z_raw[np.isfinite(filt_z_raw)]
qs = np.linspace(0.001, 0.999, 200)
norm_q = norm.ppf(qs)
axes[1].scatter(
norm_q, np.quantile(all_z, qs), s=10, alpha=0.7, label="All returns", color=COLORS["blue"]
)
axes[1].scatter(
norm_q, np.quantile(filt_z, qs), s=10, alpha=0.7, label="Jump-filtered", color=COLORS["amber"]
)
axes[1].plot([-4, 4], [-4, 4], ls="--", lw=0.8, color=COLORS["neutral"])
axes[1].set_xlim(-4, 4)
axes[1].set_ylim(-8, 8)
axes[1].set_xlabel("Theoretical quantile, N(0,1) (std. dev.)")
axes[1].set_ylabel("Empirical quantile (std. dev.)")
axes[1].set_title("Standardized returns against the normal, with and without flagged jumps")
axes[1].legend()
show_with_alt(
fig,
"Two panels. The left is an overlaid histogram, one series per symbol, of how many jumps were detected per trading day. The right is a quantile-quantile plot of standardized returns against the standard normal, with one series for all returns and another for returns with flagged jumps removed, and a dashed forty-five degree reference line.",
)
```
## 5. Jumps cluster in time of day
The opening minutes and the close concentrate price discovery: news
released overnight is impounded at the auction, and end-of-day
rebalancing flows produce sharp moves. The session-boundary fix means
the first `LEE_MYKLAND_K - 1` bars of each day (≈55 min) are not
testable, so any "first 30 min" detection here is limited; we report
the share within the testable window and look for the close-of-day
uptick that the test can see in full.
```python
tod = flagged.filter(pl.col("is_jump")).with_columns(
minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
+ pl.col("timestamp").dt.minute()
)
tod_hist = tod.group_by("minute_of_day").len().sort("minute_of_day")
all_minutes = (
flagged.with_columns(
minute_of_day=pl.col("timestamp").dt.hour().cast(pl.Int32) * 60
+ pl.col("timestamp").dt.minute()
)
.group_by("minute_of_day")
.len()
.rename({"len": "all_bars"})
.sort("minute_of_day")
)
tod_rate = tod_hist.join(all_minutes, on="minute_of_day", how="right").with_columns(
rate_per_1k=pl.col("len").fill_null(0) / pl.col("all_bars") * 1000
)
share_first_30 = tod.filter(pl.col("minute_of_day") < 9 * 60 + 30 + 30).height / tod.height
share_last_30 = tod.filter(pl.col("minute_of_day") >= 16 * 60 - 30).height / tod.height
print(f"Share of jumps in first 30 min: {share_first_30:.1%}")
print(f"Share of jumps in last 30 min: {share_last_30:.1%}")
```
```python
fig, ax = plt.subplots(figsize=(9, 3.5))
ax.bar(
tod_rate["minute_of_day"].to_numpy() / 60,
tod_rate["rate_per_1k"].to_numpy(),
width=BAR_MINUTES / 60.0,
edgecolor="none",
color=COLORS["blue"],
)
ax.set_xlabel("Hour of day (ET)")
ax.set_ylabel("Jumps per 1,000 bars")
ax.set_title("Jump rate by time of day")
ax.axvspan(9.5, 10.0, alpha=0.12, color=COLORS["neutral"])
ax.axvspan(15.5, 16.0, alpha=0.12, color=COLORS["neutral"])
show_with_alt(
fig,
"A chart of the rate at which jumps are detected against time of day across the trading session.",
)
```
## 6. Variance decomposition over the year
Stacking the continuous (bipower) and jump variance components over
time shows that jump contribution is episodic: most days have a small
or zero jump variance, punctuated by clusters around earnings,
macroeconomic releases, and stress periods.
```python
fig, ax = plt.subplots(figsize=(11, 4))
focus = daily.filter(pl.col("symbol") == SYMBOLS[0]).sort("date").to_pandas()
ax.fill_between(
focus["date"], 0, focus["bv"], alpha=0.7, color=COLORS["neutral"], label="Continuous (BV)"
)
ax.fill_between(
focus["date"],
focus["bv"],
focus["bv"] + focus["jump_var"],
alpha=0.85,
color=COLORS["amber"],
label="Jump variance (RV − BV)",
)
ax.set_xlabel("Date (2020)")
ax.set_ylabel("Daily realized variance (sum of squared log returns)")
ax.set_title(f"{SYMBOLS[0]}: daily jump and diffusive variance over time")
ax.legend()
show_with_alt(
fig,
"Two series over time for one symbol: the daily variance attributed to jumps and the daily variance attributed to continuous diffusion, on the same axis.",
)
```
## 7. Why the local-volatility adjustment matters
A naive rule that flags any bar with $|r_i / \hat\sigma_d| > 4$, where
$\hat\sigma_d$ is a single daily volatility, ignores time-of-day
volatility shape and slow regime changes. Compared with Lee–Mykland
it over-detects on calm days (the daily $\hat\sigma_d$ is small, so
even modestly elevated bars exceed the fixed $4\hat\sigma_d$
threshold) and under-detects on volatile days (the daily $\hat\sigma_d$
is inflated by the very moves we are trying to flag, so large bars
fail to clear $4\hat\sigma_d$).
```python
daily_sigma = returns.group_by(["symbol", "date"]).agg(daily_std=pl.col("log_return").std())
naive = (
returns.join(daily_sigma, on=["symbol", "date"])
.with_columns(naive_z=pl.col("log_return") / pl.col("daily_std"))
.with_columns(naive_jump=pl.col("naive_z").abs() > NAIVE_Z_THRESHOLD)
)
comparison = (
flagged.join(naive.select("symbol", "timestamp", "naive_jump"), on=["symbol", "timestamp"])
.group_by("symbol")
.agg(
lm_jumps=pl.col("is_jump").sum(),
naive_jumps=pl.col("naive_jump").sum(),
both=(pl.col("is_jump") & pl.col("naive_jump")).sum(),
lm_only=(pl.col("is_jump") & ~pl.col("naive_jump")).sum(),
naive_only=(~pl.col("is_jump") & pl.col("naive_jump")).sum(),
)
.sort("symbol")
)
```
Split each symbol's flags into three buckets: bars both rules agree are
jumps, bars only Lee–Mykland flags (the naive rule *misses* these, mostly
on volatile days when the inflated daily $\hat\sigma_d$ swallows real
jumps), and bars only the naive rule flags (it *over-fires* on calm days
when the small daily $\hat\sigma_d$ makes ordinary bars look extreme).
Substantial disagreement on either side means the fixed-threshold rule is
not a safe substitute.
```python
syms = comparison["symbol"].to_list()
x = np.arange(len(syms))
w = 0.26
fig, ax = plt.subplots(figsize=(9, 4))
ax.bar(x - w, comparison["both"].to_numpy(), w, color=COLORS["blue"], label="Both agree")
ax.bar(x, comparison["lm_only"].to_numpy(), w, color=COLORS["amber"], label="Lee–Mykland only")
ax.bar(x + w, comparison["naive_only"].to_numpy(), w, color=COLORS["copper"], label="Naive only")
ax.set_xticks(x)
ax.set_xticklabels(syms)
ax.set_xlabel("Symbol")
ax.set_ylabel("Flagged bars (count)")
ax.set_title("Bars flagged by a fixed return threshold against Lee-Mykland")
ax.legend()
show_with_alt(
fig,
"A comparison of which bars two rules flag as jumps: a fixed return threshold and the Lee-Mykland test, showing where they agree and where each fires alone.",
)
```
## 8. Materialize the jump-feature panel
The output is a per-symbol, per-day frame with the features that
downstream chapters consume directly: total jump count, signed jump
variance, and the jump share of total realized variance. These join
cleanly to the bar-level feature panels of Chapters 7–8.
```python
features = daily_jump_counts.select(
"symbol",
"date",
"jump_count",
"signed_jump_var",
pl.col("jump_var").alias("jump_variance"),
pl.col("jump_share_rv").alias("jump_share"),
pl.col("rv").alias("rv_total"),
pl.col("bv").alias("continuous_variance"),
).sort(["symbol", "date"])
features.write_parquet(OUTPUT_DIR / "daily_jump_features.parquet")
print(
f"Wrote {features.height:,} rows to {display_path(OUTPUT_DIR / 'daily_jump_features.parquet')}"
)
features.head(10)
```
## Takeaways
- Realized variance over-states diffusive risk because it absorbs
jumps; bipower variation gives a jump-robust continuous-variance
estimate and the difference times jump contribution.
- The Lee–Mykland test fires the multiple-testing-correct number of
jumps per day in expectation and adapts to time-varying volatility,
so it is comparable across regimes in a way that fixed-threshold
rules are not.
- The Lee–Mykland jump rate by time of day shows a clear close-of-day
uptick within the testable window (the first `LEE_MYKLAND_K - 1` bars
of each session are not directly testable). See §5 for the printed
per-window shares. The notebook does not condition on scheduled news
days; the daily variance decomposition in §6 shows episodic jump
contribution but is not aligned to an external news calendar.
- The daily features written above plug into Chapter 7 (jump-conditional
labels and event-time sampling) and Chapter 8 (volatility-regime
features and feature stratification).




출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.