跳至正文
返回文库全部文档

构建与调校成交笔数及成交量失衡柱

代码 《交易机器学习》

总结

本笔记通过交易数据构建成交笔数和成交量失衡柱,考察由信息驱动的抽样方法。它将手动实现的成交笔数失衡算法与库中的采样器进行核验,随后调整预期柱大小和自适应设置,比较自适应、固定阈值和滚动窗口方法。通过柱数、平均大小、Jarque–Bera 正态性检验、滞后一阶自相关和方差比统计量评估柱收益。该框架强调随时间检查采样器的阈值组成:仅观察期末预期规模的漂移,可能遗漏自适应过程中向上、向下或相互抵消的变化。

文中说明,失衡柱会累积带符号的交易活动,直到达到阈值;自适应方案则会更新预期规模和买入概率。这些更新可能造成阈值不稳定,使柱变得过大,或接近每笔交易形成一根柱。若可预测的柱大小很重要,建议使用固定采样器;研究时则建议使用缓慢自适应的采样器。结论仅适用于所抽取的 NVDA 交易流和所选参数;分布诊断本身无法证明存在盈利信号或广泛的市场规律。

核心观点

  • 成交笔数失衡会累积带符号的交易方向,直到达到预期失衡阈值。 成交量失衡将相同的抽样思路应用于带符号的交易量。 自适应预期柱大小和交易方向概率可能同时变化,掩盖阈值不稳定性。 应随时间检查阈值组成,并同时关注柱数和平均柱大小。 统计柱诊断描述的是抽样行为,无法证明存在交易优势。

标签

全文
# 16_itch_information_bars.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Information-Driven Bars: Formulas and Parameter Study
#
# **Chapter 3: Market Microstructure**
#
# **Docker image**: `ml4t`
#
# ## Purpose
#
# Two-part validation of imbalance-bar construction: (1) verify the AFML
# tick-imbalance formula matches the `ml4t.engineer.bars` library exactly on
# DataBento NVDA trades, and (2) sweep $\alpha$ and the target $E[T]$ across
# three families (alpha-based EWMA, fixed threshold, rolling window) to expose
# the parameter-instability issue that §3.4 warns about.
#
# ## Learning Objectives
#
# After completing this notebook, you will be able to:
# - State the tick-imbalance threshold formula
#   $\theta = \sum b_t$, $E[\theta_T] = E[T] \cdot |2P[b=1] - 1|$
#   and the volume-imbalance variant.
# - Compare an adaptive threshold, a fixed one and a rolling-window one on the same
#   trade stream, and recognise the two ways the adaptive scheme fails: a threshold
#   that runs away upward, and one that falls until every trade cuts a bar.
# - Diagnose which of those is happening from the bar count and the average bar size,
#   before looking at any downstream statistic.
# - Choose imbalance-bar parameters that yield well-behaved
#   Jarque-Bera / variance-ratio diagnostics.
#
# ## Book reference
#
# Section §3.4, *The Art of Sampling* — information-driven-bars subsection.
#
# ## Prerequisites
#
# - DataBento XNAS-ITCH MBO parquets at
#   `data/equities/market/microstructure/market_by_order/NVDA/`.

# %%
"""Information-Driven Bars: Formulas and Parameter Study — verifying AFML imbalance bar formulas and exploring parameter sensitivity."""

import re
from pathlib import Path

import numpy as np
import plotly.graph_objects as go
import polars as pl
from plotly.subplots import make_subplots
from scipy import stats

# Import loader for MBO data
from data import load_mbo_data
from utils.style import show_plotly_with_alt

# Polars display configuration

# %% tags=["parameters"]
MAX_DAYS = 3

# %%
# Get file paths from the canonical loader (handles legacy/new path resolution)
data_files = load_mbo_data(symbols=["NVDA"], list_files=True)
DATABENTO_DIR = data_files[0].parent if data_files else None

# %% [markdown]
# ### Load Trade Data
# Extract and filter trade records from multi-day MBO parquet files.


# %%
def load_trades(data_dir: Path, max_days: int = 10) -> tuple[pl.DataFrame | None, list[str]]:
    """Load trade data from multiple days."""
    data_files = sorted(data_dir.glob("*.parquet"))[:max_days]
    if not data_files:
        return None, []

    all_trades = []
    dates = []

    for file_path in data_files:
        # Derive the date from the trailing 8-digit token so both file layouts
        # work: Download Center `xnas-itch-YYYYMMDD.mbo.dbn.parquet` and API
        # (`mbo_download.py`) `YYYYMMDD.parquet`.
        date_str = re.search(r"(\d{8})", file_path.stem).group(1)
        trade_date = f"{date_str[:4]}-{date_str[4:6]}-{date_str[6:8]}"
        dates.append(trade_date)

        df = pl.read_parquet(file_path)
        # Prefer exchange event time; fall back across the two file layouts.
        ts_col = next(c for c in ("timestamp", "ts_event", "ts_recv") if c in df.columns)
        df = df.select([ts_col, "action", "side", "price", "size"])
        df = df.with_columns(pl.col(ts_col).cast(pl.Datetime("ns")).alias("timestamp"))
        df = df.filter(pl.col("action") == "T")
        # Regular trading hours (09:30-16:00 America/New_York). Convert the UTC
        # instant to exchange-local time so the window is correct in both EDT and
        # EST rather than admitting an hour of pre-market (as a fixed UTC window does).
        _et = (
            pl.col("timestamp").dt.replace_time_zone("UTC").dt.convert_time_zone("America/New_York")
        )
        df = df.filter(
            ((_et.dt.hour() > 9) | ((_et.dt.hour() == 9) & (_et.dt.minute() >= 30)))
            & (_et.dt.hour() < 16)
        )
        # Aggressor side for Trade (T) records: DataBento sets `side` to the trade
        # aggressor — B = buy-initiated (+1), A = sell-initiated (-1). (The
        # resting-order interpretation applies to Fill `F` records, not `T`.)
        df = df.with_columns(
            pl.when(pl.col("side") == "B")
            .then(1)
            .when(pl.col("side") == "A")
            .then(-1)
            .otherwise(0)
            .alias("side_num")
        )
        df = df.select(
            [
                "timestamp",
                pl.col("price"),
                pl.col("size").alias("volume"),
                pl.col("side_num").alias("side"),
            ]
        ).sort("timestamp")
        all_trades.append(df)

    return pl.concat(all_trades), dates


# %%
# Load data
trades, dates = load_trades(DATABENTO_DIR, max_days=MAX_DAYS)

if trades is None or len(trades) == 0:
    raise FileNotFoundError(
        "Missing DataBento MBO trade data for NVDA. "
        "Expected parquet files under data/equities/market_by_order/NVDA."
    )

trades = trades.filter(pl.col("side") != 0)
print(f"Loaded {len(trades):,} trades from {len(dates)} days")
print(f"Date range: {dates[0]} to {dates[-1]}")
print(f"Buy fraction: {(trades['side'] > 0).mean():.2%}")

# %% [markdown]
# ## 2. Formula Verification: Manual vs Library
#
# We implement tick imbalance bars manually to verify the library is correct.


# %%
def calculate_tick_imbalance_bars_manual(
    sides: np.ndarray,
    expected_t: float = 1000.0,
    alpha: float = 0.1,
    min_bars_warmup: int = 10,
) -> tuple[list[int], list[dict]]:
    """
    Manual AFML tick imbalance bars.

    θ = Σ b_t (cumulative signed ticks)
    E[θ_T] = E[T] × |2P[b=1] - 1|
    """
    n = len(sides)

    # Initialize from warmup (matches library)
    warmup_size = min(1000, n)
    p_buy = float(np.mean(sides[:warmup_size] > 0))

    bar_indices = []
    bar_info = []

    cumulative_theta = 0.0
    bar_tick_count = 0
    bar_buy_count = 0
    n_bars = 0

    for i in range(n):
        side = sides[i]
        is_buy = side > 0

        cumulative_theta += side
        bar_tick_count += 1
        if is_buy:
            bar_buy_count += 1

        # AFML threshold
        threshold = expected_t * abs(2 * p_buy - 1)

        if abs(cumulative_theta) >= threshold:
            bar_indices.append(i)
            bar_info.append(
                {
                    "bar": n_bars,
                    "ticks": bar_tick_count,
                    "theta": cumulative_theta,
                    "threshold": threshold,
                    "E[T]": expected_t,
                    "P[b=1]": p_buy,
                }
            )

            n_bars += 1

            # Update EWMA after warmup
            if n_bars > min_bars_warmup:
                expected_t = alpha * bar_tick_count + (1 - alpha) * expected_t
                bar_p_buy = bar_buy_count / bar_tick_count
                p_buy = alpha * bar_p_buy + (1 - alpha) * p_buy

            # Reset
            cumulative_theta = 0.0
            bar_tick_count = 0
            bar_buy_count = 0

    return bar_indices, bar_info


# %% [markdown]
# The verification below runs the sampler by hand on the trade signs, with a slow decay
# and a long warm-up, so that the threshold sequence can be inspected step by step
# rather than only through the bars it produced.

# %%
sides_arr = trades["side"].to_numpy()
VERIFY_ET = 1000
VERIFY_ALPHA = 0.001
VERIFY_WARMUP = 100

manual_indices, manual_info = calculate_tick_imbalance_bars_manual(
    sides_arr, expected_t=VERIFY_ET, alpha=VERIFY_ALPHA, min_bars_warmup=VERIFY_WARMUP
)
print(f"Manual TIB calculation: {len(manual_indices)} bars")

# %%
# Compare with library
from ml4t.engineer.bars import TickImbalanceBarSampler

sampler = TickImbalanceBarSampler(
    expected_ticks_per_bar=VERIFY_ET,
    alpha=VERIFY_ALPHA,
    min_bars_warmup=VERIFY_WARMUP,
)
library_bars = sampler.sample(trades)
print(f"Library TIB calculation: {len(library_bars)} bars")

# Verify match
if len(manual_indices) == len(library_bars):
    manual_thresholds = [d["threshold"] for d in manual_info]
    library_thresholds = library_bars["expected_imbalance"].to_list()
    max_diff = max(
        abs(m - lib) for m, lib in zip(manual_thresholds, library_thresholds, strict=False)
    )
    print(f"Max threshold difference: {max_diff:.6f}")
    print("[OK] Manual and library match!")
else:
    print("[FAIL] Bar counts differ")

# Show first few bars
print("\nFirst 5 bars:")
pl.DataFrame(manual_info[:5])

# %% [markdown]
# ## 3. Parameter Study Using Library
#
# Now we use the faster library implementation to study how E[T] affects properties.
#
# **Statistical Metrics Explained**:
# - **Jarque-Bera (JB)**: Tests normality. Lower = more normal (JB=0 is perfectly normal).
#   High JB indicates fat tails/skewness.
# - **Autocorrelation(1)**: Correlation of returns with 1-bar-lagged returns.
#   Should be ~0 for efficient markets.
# - **Variance Ratio(5)**: Var(5-bar returns) / (5 × Var(1-bar returns)).
#   Should be ~1 for random walk. >1 = momentum, <1 = mean reversion.

# %%
from ml4t.engineer.bars import ImbalanceBarSampler, TickImbalanceBarSampler


def compute_stats(bars: pl.DataFrame) -> dict:
    """Compute statistical properties of bar returns."""
    if len(bars) < 30:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    returns = bars["close"].pct_change().drop_nulls().to_numpy()
    returns = returns[np.isfinite(returns)]
    if len(returns) < 10:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    jb, _ = stats.jarque_bera(returns)
    ac = np.corrcoef(returns[:-1], returns[1:])[0, 1] if len(returns) > 1 else np.nan

    if len(returns) > 5:
        var_1 = np.var(returns)
        summed = np.array([np.sum(returns[i : i + 5]) for i in range(len(returns) - 4)])
        var_5 = np.var(summed) / 5
        vr = var_5 / var_1 if var_1 > 0 else np.nan
    else:
        vr = np.nan

    return {"n_bars": len(bars), "jarque_bera": jb, "autocorr_1": ac, "variance_ratio_5": vr}


# %% [markdown]
# The sweep below runs each sampler over a grid of target bar sizes. Both grids are
# passed as an expected number of trades per bar, so they are in the same units; they
# differ in range because the two samplers accumulate different quantities to reach
# their stopping threshold - trade signs in one case and signed volume in the other -
# and so need different targets to produce a comparable number of bars.
#
# Both brackets are chosen to produce a few hundred to a few thousand bars from this
# session, which is the range where the downstream diagnostics have enough bars to be
# meaningful and few enough that each holds real information.
#
# The decay rate is fixed at the slow end across the whole sweep, so what varies is the
# target and not the sampler's stability.

# %%
TIB_ET = [500, 700, 1000, 1500, 2000, 3000]
VIB_ET = [2000, 5000, 10000, 20000, 50000]
ALPHA = 0.001  # slow decay; the feedback loop above is what this damps
WARMUP = 100  # Longer warmup

print("=" * 70)
print("TICK IMBALANCE BARS (TIBs) - alpha=0.001, warmup=100")
print("=" * 70)
tib_results = []
for et in TIB_ET:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    tib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )

print("\n" + "=" * 70)
print("VOLUME IMBALANCE BARS (VIBs) - alpha=0.001, warmup=100")
print("=" * 70)
vib_results = []
for et in VIB_ET:
    bars = ImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    vib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )

# %% [markdown]
# ## 4. Visualize Results

# %%
tib_df = pl.DataFrame(tib_results)
vib_df = pl.DataFrame(vib_results)

# %%
fig = make_subplots(
    rows=2,
    cols=2,
    subplot_titles=[
        "Bar Count vs E[T]",
        "Jarque-Bera vs Bar Count",
        "Autocorrelation(1) vs Bar Count",
        "Variance Ratio(5) vs Bar Count",
    ],
    vertical_spacing=0.15,
    horizontal_spacing=0.12,
)

tib_color, vib_color = "#1e3a5f", "#c74b16"

# Panel 1: bar count vs E[T]
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["expected_t"].to_list(),
            y=df["n_bars"].to_list(),
            mode="lines+markers",
            name=name,
            line=dict(color=color),
        ),
        row=1,
        col=1,
    )
fig.update_xaxes(type="log", title_text="E[T]", row=1, col=1)
fig.update_yaxes(type="log", title_text="Number of Bars", row=1, col=1)

# Panel 2: Jarque-Bera vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["jarque_bera"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=1,
        col=2,
    )
fig.update_xaxes(type="log", title_text="Number of Bars", row=1, col=2)
fig.update_yaxes(type="log", title_text="Jarque-Bera", row=1, col=2)

# Panel 3: autocorrelation(1) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["autocorr_1"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=1,
    )
fig.add_hline(y=0, line_dash="dash", line_color="gray", row=2, col=1)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=1)
fig.update_yaxes(title_text="Autocorrelation(1)", row=2, col=1)

# Panel 4: variance ratio(5) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["variance_ratio_5"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=2,
    )
fig.add_hline(y=1, line_dash="dash", line_color="gray", row=2, col=2)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=2)
fig.update_yaxes(title_text="Variance Ratio(5)", row=2, col=2)

fig.update_layout(
    title="Bar-count and return diagnostics for tick and volume imbalance bars",
    height=650,
    legend=dict(x=0.5, y=1.02, xanchor="center", orientation="h"),
)
show_plotly_with_alt(
    fig,
    "Four panels in a two-by-two grid comparing tick imbalance bars against volume imbalance bars, with one series per bar type in each. The top-left panel plots the number of bars produced against the target bar size, both on logarithmic axes. The other three plot a diagnostic against the number of bars produced on a logarithmic horizontal axis: the Jarque-Bera statistic, the lag-one autocorrelation, and the five-period variance ratio.",
)

# %% [markdown]
# ## 5. Key Takeaways
#
# | Property | Tick imbalance bars | Volume imbalance bars |
# |----------|---------------------|-----------------------|
# | Accumulates | Trade signs, plus or minus one | Signed volume |
# | Threshold units | Trades | Shares |
# | Bars at the same target size | Many more | Far fewer |
#
# The threshold scales differ by orders of magnitude because they are in different
# units, so a value calibrated for one is meaningless for the other.
#
# ### Why the adaptive threshold can run away
#
# The threshold is set from a running estimate of two things: the expected bar size and
# the expected imbalance. Both are estimated from the bars the sampler has already cut,
# which is what makes the scheme self-referential.
#
# When order flow is persistently one-sided - which real data usually is - the estimated
# imbalance rises, the threshold rises with it, the next bar takes longer to fill, and
# the estimate rises again. The same loop runs in the other direction: an estimate that
# falls produces shorter bars, which lower the estimate further, until every trade cuts
# a bar. Both are the same feedback and the decay rate is what governs it.
#
# A slower decay and a longer warm-up damp the loop, at the cost of a sampler that is
# less adaptive. The comparison below runs three decay rates so the two failure
# directions and the working case can be seen against each other.
#
# ### Choosing a target bar size
#
# There is no optimal target. Choose it against:
# - Desired bar frequency (more bars = better normality, but more noise)
# - Trading horizon (intraday needs more bars than swing)
# - Signal strength vs statistical properties tradeoff

# %% [markdown]
# ## 6. Comparing Three Approaches: α-Based, Fixed, and Window-Based
#
# The ml4t-engineer library provides three different implementations:
#
# 1. **α-Based (AFML)**: Exponential decay for E[T] and P[b=1] - requires careful α tuning
# 2. **Fixed Threshold**: No adaptation - simplest and most predictable
# 3. **Window-Based**: Rolling window adaptation - bounded drift
#
# Let's compare them on the same data.

# %%
from ml4t.engineer.bars import (
    FixedTickImbalanceBarSampler,  # Fixed threshold
    TickImbalanceBarSampler,  # α-based
    WindowTickImbalanceBarSampler,  # Window-based
)

# Test parameters - match the handoff comparison
COMPARE_ET = 1000
COMPARE_THRESHOLD = 100  # For fixed


# Track how E[T] drifts for each method
def measure_et_drift(bars: pl.DataFrame) -> float:
    """Measure E[T] drift (last / first expected_t)."""
    if "expected_t" not in bars.columns or len(bars) == 0:
        return 1.0  # No drift for fixed or empty bars
    first = bars["expected_t"][0]
    last = bars["expected_t"][-1]
    return last / first if first > 0 else 1.0


# %%
print("=" * 70)
print("COMPARING THREE TICK IMBALANCE BAR APPROACHES")
print("=" * 70)

comparison = []

# 1. α-based with different alphas
for alpha in [0.001, 0.01, 0.1]:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=COMPARE_ET,
        alpha=alpha,
        min_bars_warmup=100,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"α={alpha}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"α-based α={alpha}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )

# %%
# 2. Fixed threshold
for thresh in [50, 100, 200]:
    bars = FixedTickImbalanceBarSampler(threshold=thresh).sample(trades)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"fixed={thresh}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": 1.0,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Fixed thresh={thresh}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift=N/A"
    )

# %%
# 3. Window-based with different tick windows
for tick_win in [2000, 5000, 10000]:
    bars = WindowTickImbalanceBarSampler(
        initial_expected_t=COMPARE_ET,
        bar_window=10,
        tick_window=tick_win,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"window={tick_win}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Window tick_win={tick_win}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )

# %%
# Summary table
compare_df = pl.DataFrame(comparison)
print("\n" + "=" * 70)
print("COMPARISON SUMMARY")
print("=" * 70)
compare_df = compare_df.with_columns(
    [
        pl.col("avg_ticks").round(0).cast(pl.Int64),
        pl.col("et_drift").round(2),
        pl.col("jb").round(1),
        pl.col("ac1").round(3),
    ]
)
print(compare_df)

# %% [markdown]
# ## 7. Recommendations
#
# | Use case | Sampler | Why |
# |----------|---------|-----|
# | Production | `FixedTickImbalanceBarSampler` | The threshold cannot drift, so bar size is predictable |
# | Research | `TickImbalanceBarSampler` with a slow decay | Follows the textbook scheme while damping the feedback loop |
#
# The table above carries two different things and they answer two different questions.
#
# The bar count and the average bar size say what the sampler *produced*. Both depend on
# the trade-sign sequence as well as on the threshold, so a bar shorter than the target
# is not on its own evidence that the threshold moved - a run of one-sided signs fills
# any threshold quickly, and the warm-up bars are short before adaptation has begun.
#
# `et_drift` gets closer, and it is worth being exact about what it is: the sampler's
# expected bar size at the last bar over its value at the first. That is one of the two
# factors the threshold is built from - the threshold is the expected bar size times the
# expected imbalance, and both adapt - so `et_drift` far from one is evidence the
# adaptation moved, while `et_drift` near one is not evidence that it did not: the two
# factors can move against each other, and an endpoint ratio says nothing about what
# happened in between.
#
# The threshold itself is what settles it, and the sampler carries the pieces. Reading
# its recorded expected bar size and expected imbalance bar by bar, rather than at the
# endpoints, is what distinguishes a threshold that grew, one that fell away, and one
# that oscillated to a similar-looking endpoint.
#
# None of this applies to the fixed sampler, which has no adapting threshold. What to
# check there is whether the bar count it produced is the one its threshold was
# calibrated for.

# %%
print("\n" + "=" * 70)
print("NOTEBOOK SUMMARY")
print("=" * 70)
print("\nTIBs:")
print(tib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nVIBs:")
print(vib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nThree-Approach Comparison:")
print(compare_df)
print("\nNotebook completed.")

```

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。