跳至正文
返回文库全部文档

失衡行情柱:公式验证与参数稳定性

笔记本 《交易机器学习》

总结

此笔记本介绍成交笔数失衡柱和成交量失衡柱,并用NVDA成交数据比较自适应、固定阈值和滚动窗口方法。成交笔数失衡会累积带符号的交易方向,直到达到阈值;预期阈值取决于每柱预期交易笔数,以及估计的买卖失衡程度。笔记本用库实现核对手工计算的成交笔数柱,并通过参数扫描考察柱数、平均柱大小、收益正态性、自相关和方差比。

一个核心要点是,自适应阈值可能变得不稳定:预期行情柱大小可能急剧上升,也可能下降到几乎每笔交易都会结束一根行情柱。行情柱数量和平均大小可提供初步线索,但文档建议直接检查预期行情柱大小和失衡程度随时间的变化,因为端点比率可能掩盖中间过程的变化。如果需要稳定可预测的行情柱大小,文档建议使用固定阈值;研究中可考虑采用较慢的自适应方式。结论取决于抽样交易流和参数选择;下游诊断需要足够多的行情柱,而且这些诊断本身不能证明某种抽样方法普遍最佳。

核心观点

  • 累计带符号的交易方向达到自适应阈值时,成交笔数失衡柱就会结束。
  • 预期阈值取决于预期柱大小和对交易方向失衡程度的估计。
  • 自适应阈值可能逐渐升高,也可能下降到行情柱极短。
  • 柱数和平均大小是有用的诊断指标,但应检查阈值各组成部分随时间的变化。
  • 固定阈值使柱大小更可预测,缓慢自适应方案则适用于研究。

标签

全文
# Information-Driven Bars: Formulas and Parameter Study


# Information-Driven Bars: Formulas and Parameter Study

**Chapter 3: Market Microstructure**

**Docker image**: `ml4t`

## Purpose

Two-part validation of imbalance-bar construction: (1) verify the AFML
tick-imbalance formula matches the `ml4t.engineer.bars` library exactly on
DataBento NVDA trades, and (2) sweep $\alpha$ and the target $E[T]$ across
three families (alpha-based EWMA, fixed threshold, rolling window) to expose
the parameter-instability issue that §3.4 warns about.

## Learning Objectives

After completing this notebook, you will be able to:
- State the tick-imbalance threshold formula
  $\theta = \sum b_t$, $E[\theta_T] = E[T] \cdot |2P[b=1] - 1|$
  and the volume-imbalance variant.
- Compare an adaptive threshold, a fixed one and a rolling-window one on the same
  trade stream, and recognise the two ways the adaptive scheme fails: a threshold
  that runs away upward, and one that falls until every trade cuts a bar.
- Diagnose which of those is happening from the bar count and the average bar size,
  before looking at any downstream statistic.
- Choose imbalance-bar parameters that yield well-behaved
  Jarque-Bera / variance-ratio diagnostics.

## Book reference

Section §3.4, *The Art of Sampling* — information-driven-bars subsection.

## Prerequisites

- DataBento XNAS-ITCH MBO parquets at
  `data/equities/market/microstructure/market_by_order/NVDA/`.

```python
"""Information-Driven Bars: Formulas and Parameter Study — verifying AFML imbalance bar formulas and exploring parameter sensitivity."""

import re
from pathlib import Path

import numpy as np
import plotly.graph_objects as go
import polars as pl
from plotly.subplots import make_subplots
from scipy import stats

# Import loader for MBO data
from data import load_mbo_data
from utils.style import show_plotly_with_alt

# Polars display configuration
```

```python
MAX_DAYS = 3
```

```python
# Get file paths from the canonical loader (handles legacy/new path resolution)
data_files = load_mbo_data(symbols=["NVDA"], list_files=True)
DATABENTO_DIR = data_files[0].parent if data_files else None
```

### Load Trade Data
Extract and filter trade records from multi-day MBO parquet files.

```python
def load_trades(data_dir: Path, max_days: int = 10) -> tuple[pl.DataFrame | None, list[str]]:
    """Load trade data from multiple days."""
    data_files = sorted(data_dir.glob("*.parquet"))[:max_days]
    if not data_files:
        return None, []

    all_trades = []
    dates = []

    for file_path in data_files:
        # Derive the date from the trailing 8-digit token so both file layouts
        # work: Download Center `xnas-itch-YYYYMMDD.mbo.dbn.parquet` and API
        # (`mbo_download.py`) `YYYYMMDD.parquet`.
        date_str = re.search(r"(\d{8})", file_path.stem).group(1)
        trade_date = f"{date_str[:4]}-{date_str[4:6]}-{date_str[6:8]}"
        dates.append(trade_date)

        df = pl.read_parquet(file_path)
        # Prefer exchange event time; fall back across the two file layouts.
        ts_col = next(c for c in ("timestamp", "ts_event", "ts_recv") if c in df.columns)
        df = df.select([ts_col, "action", "side", "price", "size"])
        df = df.with_columns(pl.col(ts_col).cast(pl.Datetime("ns")).alias("timestamp"))
        df = df.filter(pl.col("action") == "T")
        # Regular trading hours (09:30-16:00 America/New_York). Convert the UTC
        # instant to exchange-local time so the window is correct in both EDT and
        # EST rather than admitting an hour of pre-market (as a fixed UTC window does).
        _et = (
            pl.col("timestamp").dt.replace_time_zone("UTC").dt.convert_time_zone("America/New_York")
        )
        df = df.filter(
            ((_et.dt.hour() > 9) | ((_et.dt.hour() == 9) & (_et.dt.minute() >= 30)))
            & (_et.dt.hour() < 16)
        )
        # Aggressor side for Trade (T) records: DataBento sets `side` to the trade
        # aggressor — B = buy-initiated (+1), A = sell-initiated (-1). (The
        # resting-order interpretation applies to Fill `F` records, not `T`.)
        df = df.with_columns(
            pl.when(pl.col("side") == "B")
            .then(1)
            .when(pl.col("side") == "A")
            .then(-1)
            .otherwise(0)
            .alias("side_num")
        )
        df = df.select(
            [
                "timestamp",
                pl.col("price"),
                pl.col("size").alias("volume"),
                pl.col("side_num").alias("side"),
            ]
        ).sort("timestamp")
        all_trades.append(df)

    return pl.concat(all_trades), dates
```

```python
# Load data
trades, dates = load_trades(DATABENTO_DIR, max_days=MAX_DAYS)

if trades is None or len(trades) == 0:
    raise FileNotFoundError(
        "Missing DataBento MBO trade data for NVDA. "
        "Expected parquet files under data/equities/market_by_order/NVDA."
    )

trades = trades.filter(pl.col("side") != 0)
print(f"Loaded {len(trades):,} trades from {len(dates)} days")
print(f"Date range: {dates[0]} to {dates[-1]}")
print(f"Buy fraction: {(trades['side'] > 0).mean():.2%}")
```

## 2. Formula Verification: Manual vs Library

We implement tick imbalance bars manually to verify the library is correct.

```python
def calculate_tick_imbalance_bars_manual(
    sides: np.ndarray,
    expected_t: float = 1000.0,
    alpha: float = 0.1,
    min_bars_warmup: int = 10,
) -> tuple[list[int], list[dict]]:
    """
    Manual AFML tick imbalance bars.

    θ = Σ b_t (cumulative signed ticks)
    E[θ_T] = E[T] × |2P[b=1] - 1|
    """
    n = len(sides)

    # Initialize from warmup (matches library)
    warmup_size = min(1000, n)
    p_buy = float(np.mean(sides[:warmup_size] > 0))

    bar_indices = []
    bar_info = []

    cumulative_theta = 0.0
    bar_tick_count = 0
    bar_buy_count = 0
    n_bars = 0

    for i in range(n):
        side = sides[i]
        is_buy = side > 0

        cumulative_theta += side
        bar_tick_count += 1
        if is_buy:
            bar_buy_count += 1

        # AFML threshold
        threshold = expected_t * abs(2 * p_buy - 1)

        if abs(cumulative_theta) >= threshold:
            bar_indices.append(i)
            bar_info.append(
                {
                    "bar": n_bars,
                    "ticks": bar_tick_count,
                    "theta": cumulative_theta,
                    "threshold": threshold,
                    "E[T]": expected_t,
                    "P[b=1]": p_buy,
                }
            )

            n_bars += 1

            # Update EWMA after warmup
            if n_bars > min_bars_warmup:
                expected_t = alpha * bar_tick_count + (1 - alpha) * expected_t
                bar_p_buy = bar_buy_count / bar_tick_count
                p_buy = alpha * bar_p_buy + (1 - alpha) * p_buy

            # Reset
            cumulative_theta = 0.0
            bar_tick_count = 0
            bar_buy_count = 0

    return bar_indices, bar_info
```

The verification below runs the sampler by hand on the trade signs, with a slow decay
and a long warm-up, so that the threshold sequence can be inspected step by step
rather than only through the bars it produced.

```python
sides_arr = trades["side"].to_numpy()
VERIFY_ET = 1000
VERIFY_ALPHA = 0.001
VERIFY_WARMUP = 100

manual_indices, manual_info = calculate_tick_imbalance_bars_manual(
    sides_arr, expected_t=VERIFY_ET, alpha=VERIFY_ALPHA, min_bars_warmup=VERIFY_WARMUP
)
print(f"Manual TIB calculation: {len(manual_indices)} bars")
```

```python
# Compare with library
from ml4t.engineer.bars import TickImbalanceBarSampler

sampler = TickImbalanceBarSampler(
    expected_ticks_per_bar=VERIFY_ET,
    alpha=VERIFY_ALPHA,
    min_bars_warmup=VERIFY_WARMUP,
)
library_bars = sampler.sample(trades)
print(f"Library TIB calculation: {len(library_bars)} bars")

# Verify match
if len(manual_indices) == len(library_bars):
    manual_thresholds = [d["threshold"] for d in manual_info]
    library_thresholds = library_bars["expected_imbalance"].to_list()
    max_diff = max(
        abs(m - lib) for m, lib in zip(manual_thresholds, library_thresholds, strict=False)
    )
    print(f"Max threshold difference: {max_diff:.6f}")
    print("[OK] Manual and library match!")
else:
    print("[FAIL] Bar counts differ")

# Show first few bars
print("\nFirst 5 bars:")
pl.DataFrame(manual_info[:5])
```

## 3. Parameter Study Using Library

Now we use the faster library implementation to study how E[T] affects properties.

**Statistical Metrics Explained**:
- **Jarque-Bera (JB)**: Tests normality. Lower = more normal (JB=0 is perfectly normal).
  High JB indicates fat tails/skewness.
- **Autocorrelation(1)**: Correlation of returns with 1-bar-lagged returns.
  Should be ~0 for efficient markets.
- **Variance Ratio(5)**: Var(5-bar returns) / (5 × Var(1-bar returns)).
  Should be ~1 for random walk. >1 = momentum, <1 = mean reversion.

```python
from ml4t.engineer.bars import ImbalanceBarSampler, TickImbalanceBarSampler


def compute_stats(bars: pl.DataFrame) -> dict:
    """Compute statistical properties of bar returns."""
    if len(bars) < 30:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    returns = bars["close"].pct_change().drop_nulls().to_numpy()
    returns = returns[np.isfinite(returns)]
    if len(returns) < 10:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    jb, _ = stats.jarque_bera(returns)
    ac = np.corrcoef(returns[:-1], returns[1:])[0, 1] if len(returns) > 1 else np.nan

    if len(returns) > 5:
        var_1 = np.var(returns)
        summed = np.array([np.sum(returns[i : i + 5]) for i in range(len(returns) - 4)])
        var_5 = np.var(summed) / 5
        vr = var_5 / var_1 if var_1 > 0 else np.nan
    else:
        vr = np.nan

    return {"n_bars": len(bars), "jarque_bera": jb, "autocorr_1": ac, "variance_ratio_5": vr}
```

The sweep below runs each sampler over a grid of target bar sizes. Both grids are
passed as an expected number of trades per bar, so they are in the same units; they
differ in range because the two samplers accumulate different quantities to reach
their stopping threshold - trade signs in one case and signed volume in the other -
and so need different targets to produce a comparable number of bars.

Both brackets are chosen to produce a few hundred to a few thousand bars from this
session, which is the range where the downstream diagnostics have enough bars to be
meaningful and few enough that each holds real information.

The decay rate is fixed at the slow end across the whole sweep, so what varies is the
target and not the sampler's stability.

```python
TIB_ET = [500, 700, 1000, 1500, 2000, 3000]
VIB_ET = [2000, 5000, 10000, 20000, 50000]
ALPHA = 0.001  # slow decay; the feedback loop above is what this damps
WARMUP = 100  # Longer warmup

print("=" * 70)
print("TICK IMBALANCE BARS (TIBs) - alpha=0.001, warmup=100")
print("=" * 70)
tib_results = []
for et in TIB_ET:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    tib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )

print("\n" + "=" * 70)
print("VOLUME IMBALANCE BARS (VIBs) - alpha=0.001, warmup=100")
print("=" * 70)
vib_results = []
for et in VIB_ET:
    bars = ImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    vib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )
```

## 4. Visualize Results

```python
tib_df = pl.DataFrame(tib_results)
vib_df = pl.DataFrame(vib_results)
```

```python
fig = make_subplots(
    rows=2,
    cols=2,
    subplot_titles=[
        "Bar Count vs E[T]",
        "Jarque-Bera vs Bar Count",
        "Autocorrelation(1) vs Bar Count",
        "Variance Ratio(5) vs Bar Count",
    ],
    vertical_spacing=0.15,
    horizontal_spacing=0.12,
)

tib_color, vib_color = "#1e3a5f", "#c74b16"

# Panel 1: bar count vs E[T]
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["expected_t"].to_list(),
            y=df["n_bars"].to_list(),
            mode="lines+markers",
            name=name,
            line=dict(color=color),
        ),
        row=1,
        col=1,
    )
fig.update_xaxes(type="log", title_text="E[T]", row=1, col=1)
fig.update_yaxes(type="log", title_text="Number of Bars", row=1, col=1)

# Panel 2: Jarque-Bera vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["jarque_bera"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=1,
        col=2,
    )
fig.update_xaxes(type="log", title_text="Number of Bars", row=1, col=2)
fig.update_yaxes(type="log", title_text="Jarque-Bera", row=1, col=2)

# Panel 3: autocorrelation(1) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["autocorr_1"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=1,
    )
fig.add_hline(y=0, line_dash="dash", line_color="gray", row=2, col=1)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=1)
fig.update_yaxes(title_text="Autocorrelation(1)", row=2, col=1)

# Panel 4: variance ratio(5) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["variance_ratio_5"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=2,
    )
fig.add_hline(y=1, line_dash="dash", line_color="gray", row=2, col=2)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=2)
fig.update_yaxes(title_text="Variance Ratio(5)", row=2, col=2)

fig.update_layout(
    title="Bar-count and return diagnostics for tick and volume imbalance bars",
    height=650,
    legend=dict(x=0.5, y=1.02, xanchor="center", orientation="h"),
)
show_plotly_with_alt(
    fig,
    "Four panels in a two-by-two grid comparing tick imbalance bars against volume imbalance bars, with one series per bar type in each. The top-left panel plots the number of bars produced against the target bar size, both on logarithmic axes. The other three plot a diagnostic against the number of bars produced on a logarithmic horizontal axis: the Jarque-Bera statistic, the lag-one autocorrelation, and the five-period variance ratio.",
)
```

## 5. Key Takeaways

| Property | Tick imbalance bars | Volume imbalance bars |
|----------|---------------------|-----------------------|
| Accumulates | Trade signs, plus or minus one | Signed volume |
| Threshold units | Trades | Shares |
| Bars at the same target size | Many more | Far fewer |

The threshold scales differ by orders of magnitude because they are in different
units, so a value calibrated for one is meaningless for the other.

### Why the adaptive threshold can run away

The threshold is set from a running estimate of two things: the expected bar size and
the expected imbalance. Both are estimated from the bars the sampler has already cut,
which is what makes the scheme self-referential.

When order flow is persistently one-sided - which real data usually is - the estimated
imbalance rises, the threshold rises with it, the next bar takes longer to fill, and
the estimate rises again. The same loop runs in the other direction: an estimate that
falls produces shorter bars, which lower the estimate further, until every trade cuts
a bar. Both are the same feedback and the decay rate is what governs it.

A slower decay and a longer warm-up damp the loop, at the cost of a sampler that is
less adaptive. The comparison below runs three decay rates so the two failure
directions and the working case can be seen against each other.

### Choosing a target bar size

There is no optimal target. Choose it against:
- Desired bar frequency (more bars = better normality, but more noise)
- Trading horizon (intraday needs more bars than swing)
- Signal strength vs statistical properties tradeoff

## 6. Comparing Three Approaches: α-Based, Fixed, and Window-Based

The ml4t-engineer library provides three different implementations:

1. **α-Based (AFML)**: Exponential decay for E[T] and P[b=1] - requires careful α tuning
2. **Fixed Threshold**: No adaptation - simplest and most predictable
3. **Window-Based**: Rolling window adaptation - bounded drift

Let's compare them on the same data.

```python
from ml4t.engineer.bars import (
    FixedTickImbalanceBarSampler,  # Fixed threshold
    TickImbalanceBarSampler,  # α-based
    WindowTickImbalanceBarSampler,  # Window-based
)

# Test parameters - match the handoff comparison
COMPARE_ET = 1000
COMPARE_THRESHOLD = 100  # For fixed


# Track how E[T] drifts for each method
def measure_et_drift(bars: pl.DataFrame) -> float:
    """Measure E[T] drift (last / first expected_t)."""
    if "expected_t" not in bars.columns or len(bars) == 0:
        return 1.0  # No drift for fixed or empty bars
    first = bars["expected_t"][0]
    last = bars["expected_t"][-1]
    return last / first if first > 0 else 1.0
```

```python
print("=" * 70)
print("COMPARING THREE TICK IMBALANCE BAR APPROACHES")
print("=" * 70)

comparison = []

# 1. α-based with different alphas
for alpha in [0.001, 0.01, 0.1]:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=COMPARE_ET,
        alpha=alpha,
        min_bars_warmup=100,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"α={alpha}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"α-based α={alpha}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )
```

```python
# 2. Fixed threshold
for thresh in [50, 100, 200]:
    bars = FixedTickImbalanceBarSampler(threshold=thresh).sample(trades)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"fixed={thresh}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": 1.0,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Fixed thresh={thresh}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift=N/A"
    )
```

```python
# 3. Window-based with different tick windows
for tick_win in [2000, 5000, 10000]:
    bars = WindowTickImbalanceBarSampler(
        initial_expected_t=COMPARE_ET,
        bar_window=10,
        tick_window=tick_win,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"window={tick_win}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Window tick_win={tick_win}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )
```

```python
# Summary table
compare_df = pl.DataFrame(comparison)
print("\n" + "=" * 70)
print("COMPARISON SUMMARY")
print("=" * 70)
compare_df = compare_df.with_columns(
    [
        pl.col("avg_ticks").round(0).cast(pl.Int64),
        pl.col("et_drift").round(2),
        pl.col("jb").round(1),
        pl.col("ac1").round(3),
    ]
)
print(compare_df)
```

## 7. Recommendations

| Use case | Sampler | Why |
|----------|---------|-----|
| Production | `FixedTickImbalanceBarSampler` | The threshold cannot drift, so bar size is predictable |
| Research | `TickImbalanceBarSampler` with a slow decay | Follows the textbook scheme while damping the feedback loop |

The table above carries two different things and they answer two different questions.

The bar count and the average bar size say what the sampler *produced*. Both depend on
the trade-sign sequence as well as on the threshold, so a bar shorter than the target
is not on its own evidence that the threshold moved - a run of one-sided signs fills
any threshold quickly, and the warm-up bars are short before adaptation has begun.

`et_drift` gets closer, and it is worth being exact about what it is: the sampler's
expected bar size at the last bar over its value at the first. That is one of the two
factors the threshold is built from - the threshold is the expected bar size times the
expected imbalance, and both adapt - so `et_drift` far from one is evidence the
adaptation moved, while `et_drift` near one is not evidence that it did not: the two
factors can move against each other, and an endpoint ratio says nothing about what
happened in between.

The threshold itself is what settles it, and the sampler carries the pieces. Reading
its recorded expected bar size and expected imbalance bar by bar, rather than at the
endpoints, is what distinguishes a threshold that grew, one that fell away, and one
that oscillated to a similar-looking endpoint.

None of this applies to the fixed sampler, which has no adapting threshold. What to
check there is whether the bar count it produced is the one its threshold was
calibrated for.

```python
print("\n" + "=" * 70)
print("NOTEBOOK SUMMARY")
print("=" * 70)
print("\nTIBs:")
print(tib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nVIBs:")
print(vib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nThree-Approach Comparison:")
print(compare_df)
print("\nNotebook completed.")
```
![notebook output](figures/p1_1.png)

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。