본문으로 건너뛰기
라이브러리 문서 전체

불균형 바: 수식 검증과 매개변수 안정성

노트북 Machine Learning for Trading

요약

이 노트북은 틱 불균형 바와 거래량 불균형 바를 설명하고, NVDA 거래 데이터를 사용해 적응형, 고정 임계값, 이동 창 방식을 비교합니다. 틱 불균형 바는 부호가 있는 거래 방향을 임계값에 도달할 때까지 누적합니다. 예상 임계값은 바당 예상 거래 수와 매수 및 매도 불균형 추정치에 따라 달라집니다. 직접 계산한 틱 바를 라이브러리 구현과 대조하며, 매개변수 스윕으로 바 개수, 평균 바 크기, 수익률 정규성, 자기상관, 분산비를 살펴봅니다.

핵심은 적응형 임계값이 불안정해질 수 있다는 점입니다. 예상 바 크기가 급격히 커지거나 줄어들어 거의 모든 거래가 바를 닫을 수도 있습니다. 바 개수와 평균 크기는 초기 단서를 제공하지만, 양 끝의 비율은 그 사이의 변화를 숨길 수 있으므로 시간에 따라 변하는 예상 바 크기와 불균형을 직접 살펴보라고 권합니다. 바 크기를 예측 가능하게 유지해야 한다면 고정 임계값을, 리서치에는 적응 속도를 느리게 설정할 것을 권합니다. 결과는 표본 거래 흐름과 매개변수 선택에 따라 달라집니다. 후속 진단에는 충분한 바가 필요하며, 이 결과만으로 보편적으로 가장 좋은 샘플링 방법을 입증할 수는 없습니다.

핵심 아이디어

  • 누적된 거래 방향의 부호가 적응형 임계값에 도달하면 틱 불균형 바가 닫힙니다.
  • 예상 임계값은 바당 예상 거래 수와 거래 부호 불균형 추정치에 따라 달라집니다.
  • 적응형 임계값은 시간이 지나며 상승하거나 하락해 바가 매우 짧아질 수 있습니다.
  • 바 개수와 평균 크기는 유용한 진단 지표지만, 임계값 구성 요소는 시간에 따라 점검해야 합니다.
  • 고정 임계값은 바 크기를 더 예측하기 쉽게 하고, 느린 적응 방식은 리서치에 유용합니다.

태그

전문
# Information-Driven Bars: Formulas and Parameter Study


# Information-Driven Bars: Formulas and Parameter Study

**Chapter 3: Market Microstructure**

**Docker image**: `ml4t`

## Purpose

Two-part validation of imbalance-bar construction: (1) verify the AFML
tick-imbalance formula matches the `ml4t.engineer.bars` library exactly on
DataBento NVDA trades, and (2) sweep $\alpha$ and the target $E[T]$ across
three families (alpha-based EWMA, fixed threshold, rolling window) to expose
the parameter-instability issue that §3.4 warns about.

## Learning Objectives

After completing this notebook, you will be able to:
- State the tick-imbalance threshold formula
  $\theta = \sum b_t$, $E[\theta_T] = E[T] \cdot |2P[b=1] - 1|$
  and the volume-imbalance variant.
- Compare an adaptive threshold, a fixed one and a rolling-window one on the same
  trade stream, and recognise the two ways the adaptive scheme fails: a threshold
  that runs away upward, and one that falls until every trade cuts a bar.
- Diagnose which of those is happening from the bar count and the average bar size,
  before looking at any downstream statistic.
- Choose imbalance-bar parameters that yield well-behaved
  Jarque-Bera / variance-ratio diagnostics.

## Book reference

Section §3.4, *The Art of Sampling* — information-driven-bars subsection.

## Prerequisites

- DataBento XNAS-ITCH MBO parquets at
  `data/equities/market/microstructure/market_by_order/NVDA/`.

```python
"""Information-Driven Bars: Formulas and Parameter Study — verifying AFML imbalance bar formulas and exploring parameter sensitivity."""

import re
from pathlib import Path

import numpy as np
import plotly.graph_objects as go
import polars as pl
from plotly.subplots import make_subplots
from scipy import stats

# Import loader for MBO data
from data import load_mbo_data
from utils.style import show_plotly_with_alt

# Polars display configuration
```

```python
MAX_DAYS = 3
```

```python
# Get file paths from the canonical loader (handles legacy/new path resolution)
data_files = load_mbo_data(symbols=["NVDA"], list_files=True)
DATABENTO_DIR = data_files[0].parent if data_files else None
```

### Load Trade Data
Extract and filter trade records from multi-day MBO parquet files.

```python
def load_trades(data_dir: Path, max_days: int = 10) -> tuple[pl.DataFrame | None, list[str]]:
    """Load trade data from multiple days."""
    data_files = sorted(data_dir.glob("*.parquet"))[:max_days]
    if not data_files:
        return None, []

    all_trades = []
    dates = []

    for file_path in data_files:
        # Derive the date from the trailing 8-digit token so both file layouts
        # work: Download Center `xnas-itch-YYYYMMDD.mbo.dbn.parquet` and API
        # (`mbo_download.py`) `YYYYMMDD.parquet`.
        date_str = re.search(r"(\d{8})", file_path.stem).group(1)
        trade_date = f"{date_str[:4]}-{date_str[4:6]}-{date_str[6:8]}"
        dates.append(trade_date)

        df = pl.read_parquet(file_path)
        # Prefer exchange event time; fall back across the two file layouts.
        ts_col = next(c for c in ("timestamp", "ts_event", "ts_recv") if c in df.columns)
        df = df.select([ts_col, "action", "side", "price", "size"])
        df = df.with_columns(pl.col(ts_col).cast(pl.Datetime("ns")).alias("timestamp"))
        df = df.filter(pl.col("action") == "T")
        # Regular trading hours (09:30-16:00 America/New_York). Convert the UTC
        # instant to exchange-local time so the window is correct in both EDT and
        # EST rather than admitting an hour of pre-market (as a fixed UTC window does).
        _et = (
            pl.col("timestamp").dt.replace_time_zone("UTC").dt.convert_time_zone("America/New_York")
        )
        df = df.filter(
            ((_et.dt.hour() > 9) | ((_et.dt.hour() == 9) & (_et.dt.minute() >= 30)))
            & (_et.dt.hour() < 16)
        )
        # Aggressor side for Trade (T) records: DataBento sets `side` to the trade
        # aggressor — B = buy-initiated (+1), A = sell-initiated (-1). (The
        # resting-order interpretation applies to Fill `F` records, not `T`.)
        df = df.with_columns(
            pl.when(pl.col("side") == "B")
            .then(1)
            .when(pl.col("side") == "A")
            .then(-1)
            .otherwise(0)
            .alias("side_num")
        )
        df = df.select(
            [
                "timestamp",
                pl.col("price"),
                pl.col("size").alias("volume"),
                pl.col("side_num").alias("side"),
            ]
        ).sort("timestamp")
        all_trades.append(df)

    return pl.concat(all_trades), dates
```

```python
# Load data
trades, dates = load_trades(DATABENTO_DIR, max_days=MAX_DAYS)

if trades is None or len(trades) == 0:
    raise FileNotFoundError(
        "Missing DataBento MBO trade data for NVDA. "
        "Expected parquet files under data/equities/market_by_order/NVDA."
    )

trades = trades.filter(pl.col("side") != 0)
print(f"Loaded {len(trades):,} trades from {len(dates)} days")
print(f"Date range: {dates[0]} to {dates[-1]}")
print(f"Buy fraction: {(trades['side'] > 0).mean():.2%}")
```

## 2. Formula Verification: Manual vs Library

We implement tick imbalance bars manually to verify the library is correct.

```python
def calculate_tick_imbalance_bars_manual(
    sides: np.ndarray,
    expected_t: float = 1000.0,
    alpha: float = 0.1,
    min_bars_warmup: int = 10,
) -> tuple[list[int], list[dict]]:
    """
    Manual AFML tick imbalance bars.

    θ = Σ b_t (cumulative signed ticks)
    E[θ_T] = E[T] × |2P[b=1] - 1|
    """
    n = len(sides)

    # Initialize from warmup (matches library)
    warmup_size = min(1000, n)
    p_buy = float(np.mean(sides[:warmup_size] > 0))

    bar_indices = []
    bar_info = []

    cumulative_theta = 0.0
    bar_tick_count = 0
    bar_buy_count = 0
    n_bars = 0

    for i in range(n):
        side = sides[i]
        is_buy = side > 0

        cumulative_theta += side
        bar_tick_count += 1
        if is_buy:
            bar_buy_count += 1

        # AFML threshold
        threshold = expected_t * abs(2 * p_buy - 1)

        if abs(cumulative_theta) >= threshold:
            bar_indices.append(i)
            bar_info.append(
                {
                    "bar": n_bars,
                    "ticks": bar_tick_count,
                    "theta": cumulative_theta,
                    "threshold": threshold,
                    "E[T]": expected_t,
                    "P[b=1]": p_buy,
                }
            )

            n_bars += 1

            # Update EWMA after warmup
            if n_bars > min_bars_warmup:
                expected_t = alpha * bar_tick_count + (1 - alpha) * expected_t
                bar_p_buy = bar_buy_count / bar_tick_count
                p_buy = alpha * bar_p_buy + (1 - alpha) * p_buy

            # Reset
            cumulative_theta = 0.0
            bar_tick_count = 0
            bar_buy_count = 0

    return bar_indices, bar_info
```

The verification below runs the sampler by hand on the trade signs, with a slow decay
and a long warm-up, so that the threshold sequence can be inspected step by step
rather than only through the bars it produced.

```python
sides_arr = trades["side"].to_numpy()
VERIFY_ET = 1000
VERIFY_ALPHA = 0.001
VERIFY_WARMUP = 100

manual_indices, manual_info = calculate_tick_imbalance_bars_manual(
    sides_arr, expected_t=VERIFY_ET, alpha=VERIFY_ALPHA, min_bars_warmup=VERIFY_WARMUP
)
print(f"Manual TIB calculation: {len(manual_indices)} bars")
```

```python
# Compare with library
from ml4t.engineer.bars import TickImbalanceBarSampler

sampler = TickImbalanceBarSampler(
    expected_ticks_per_bar=VERIFY_ET,
    alpha=VERIFY_ALPHA,
    min_bars_warmup=VERIFY_WARMUP,
)
library_bars = sampler.sample(trades)
print(f"Library TIB calculation: {len(library_bars)} bars")

# Verify match
if len(manual_indices) == len(library_bars):
    manual_thresholds = [d["threshold"] for d in manual_info]
    library_thresholds = library_bars["expected_imbalance"].to_list()
    max_diff = max(
        abs(m - lib) for m, lib in zip(manual_thresholds, library_thresholds, strict=False)
    )
    print(f"Max threshold difference: {max_diff:.6f}")
    print("[OK] Manual and library match!")
else:
    print("[FAIL] Bar counts differ")

# Show first few bars
print("\nFirst 5 bars:")
pl.DataFrame(manual_info[:5])
```

## 3. Parameter Study Using Library

Now we use the faster library implementation to study how E[T] affects properties.

**Statistical Metrics Explained**:
- **Jarque-Bera (JB)**: Tests normality. Lower = more normal (JB=0 is perfectly normal).
  High JB indicates fat tails/skewness.
- **Autocorrelation(1)**: Correlation of returns with 1-bar-lagged returns.
  Should be ~0 for efficient markets.
- **Variance Ratio(5)**: Var(5-bar returns) / (5 × Var(1-bar returns)).
  Should be ~1 for random walk. >1 = momentum, <1 = mean reversion.

```python
from ml4t.engineer.bars import ImbalanceBarSampler, TickImbalanceBarSampler


def compute_stats(bars: pl.DataFrame) -> dict:
    """Compute statistical properties of bar returns."""
    if len(bars) < 30:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    returns = bars["close"].pct_change().drop_nulls().to_numpy()
    returns = returns[np.isfinite(returns)]
    if len(returns) < 10:
        return {
            "n_bars": len(bars),
            "jarque_bera": np.nan,
            "autocorr_1": np.nan,
            "variance_ratio_5": np.nan,
        }

    jb, _ = stats.jarque_bera(returns)
    ac = np.corrcoef(returns[:-1], returns[1:])[0, 1] if len(returns) > 1 else np.nan

    if len(returns) > 5:
        var_1 = np.var(returns)
        summed = np.array([np.sum(returns[i : i + 5]) for i in range(len(returns) - 4)])
        var_5 = np.var(summed) / 5
        vr = var_5 / var_1 if var_1 > 0 else np.nan
    else:
        vr = np.nan

    return {"n_bars": len(bars), "jarque_bera": jb, "autocorr_1": ac, "variance_ratio_5": vr}
```

The sweep below runs each sampler over a grid of target bar sizes. Both grids are
passed as an expected number of trades per bar, so they are in the same units; they
differ in range because the two samplers accumulate different quantities to reach
their stopping threshold - trade signs in one case and signed volume in the other -
and so need different targets to produce a comparable number of bars.

Both brackets are chosen to produce a few hundred to a few thousand bars from this
session, which is the range where the downstream diagnostics have enough bars to be
meaningful and few enough that each holds real information.

The decay rate is fixed at the slow end across the whole sweep, so what varies is the
target and not the sampler's stability.

```python
TIB_ET = [500, 700, 1000, 1500, 2000, 3000]
VIB_ET = [2000, 5000, 10000, 20000, 50000]
ALPHA = 0.001  # slow decay; the feedback loop above is what this damps
WARMUP = 100  # Longer warmup

print("=" * 70)
print("TICK IMBALANCE BARS (TIBs) - alpha=0.001, warmup=100")
print("=" * 70)
tib_results = []
for et in TIB_ET:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    tib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )

print("\n" + "=" * 70)
print("VOLUME IMBALANCE BARS (VIBs) - alpha=0.001, warmup=100")
print("=" * 70)
vib_results = []
for et in VIB_ET:
    bars = ImbalanceBarSampler(
        expected_ticks_per_bar=et, alpha=ALPHA, min_bars_warmup=WARMUP
    ).sample(trades)
    s = compute_stats(bars)
    s["expected_t"] = et
    vib_results.append(s)
    print(
        f"E[T]={et:>6}: {s['n_bars']:>5} bars, JB={s['jarque_bera']:>8.1f}, "
        f"AC(1)={s['autocorr_1']:>6.3f}, VR(5)={s['variance_ratio_5']:>5.2f}"
    )
```

## 4. Visualize Results

```python
tib_df = pl.DataFrame(tib_results)
vib_df = pl.DataFrame(vib_results)
```

```python
fig = make_subplots(
    rows=2,
    cols=2,
    subplot_titles=[
        "Bar Count vs E[T]",
        "Jarque-Bera vs Bar Count",
        "Autocorrelation(1) vs Bar Count",
        "Variance Ratio(5) vs Bar Count",
    ],
    vertical_spacing=0.15,
    horizontal_spacing=0.12,
)

tib_color, vib_color = "#1e3a5f", "#c74b16"

# Panel 1: bar count vs E[T]
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["expected_t"].to_list(),
            y=df["n_bars"].to_list(),
            mode="lines+markers",
            name=name,
            line=dict(color=color),
        ),
        row=1,
        col=1,
    )
fig.update_xaxes(type="log", title_text="E[T]", row=1, col=1)
fig.update_yaxes(type="log", title_text="Number of Bars", row=1, col=1)

# Panel 2: Jarque-Bera vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["jarque_bera"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=1,
        col=2,
    )
fig.update_xaxes(type="log", title_text="Number of Bars", row=1, col=2)
fig.update_yaxes(type="log", title_text="Jarque-Bera", row=1, col=2)

# Panel 3: autocorrelation(1) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["autocorr_1"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=1,
    )
fig.add_hline(y=0, line_dash="dash", line_color="gray", row=2, col=1)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=1)
fig.update_yaxes(title_text="Autocorrelation(1)", row=2, col=1)

# Panel 4: variance ratio(5) vs bar count
for name, color, df in [("TIB", tib_color, tib_df), ("VIB", vib_color, vib_df)]:
    fig.add_trace(
        go.Scatter(
            x=df["n_bars"].to_list(),
            y=df["variance_ratio_5"].to_list(),
            mode="markers",
            marker=dict(color=color, size=10),
            showlegend=False,
        ),
        row=2,
        col=2,
    )
fig.add_hline(y=1, line_dash="dash", line_color="gray", row=2, col=2)
fig.update_xaxes(type="log", title_text="Number of Bars", row=2, col=2)
fig.update_yaxes(title_text="Variance Ratio(5)", row=2, col=2)

fig.update_layout(
    title="Bar-count and return diagnostics for tick and volume imbalance bars",
    height=650,
    legend=dict(x=0.5, y=1.02, xanchor="center", orientation="h"),
)
show_plotly_with_alt(
    fig,
    "Four panels in a two-by-two grid comparing tick imbalance bars against volume imbalance bars, with one series per bar type in each. The top-left panel plots the number of bars produced against the target bar size, both on logarithmic axes. The other three plot a diagnostic against the number of bars produced on a logarithmic horizontal axis: the Jarque-Bera statistic, the lag-one autocorrelation, and the five-period variance ratio.",
)
```

## 5. Key Takeaways

| Property | Tick imbalance bars | Volume imbalance bars |
|----------|---------------------|-----------------------|
| Accumulates | Trade signs, plus or minus one | Signed volume |
| Threshold units | Trades | Shares |
| Bars at the same target size | Many more | Far fewer |

The threshold scales differ by orders of magnitude because they are in different
units, so a value calibrated for one is meaningless for the other.

### Why the adaptive threshold can run away

The threshold is set from a running estimate of two things: the expected bar size and
the expected imbalance. Both are estimated from the bars the sampler has already cut,
which is what makes the scheme self-referential.

When order flow is persistently one-sided - which real data usually is - the estimated
imbalance rises, the threshold rises with it, the next bar takes longer to fill, and
the estimate rises again. The same loop runs in the other direction: an estimate that
falls produces shorter bars, which lower the estimate further, until every trade cuts
a bar. Both are the same feedback and the decay rate is what governs it.

A slower decay and a longer warm-up damp the loop, at the cost of a sampler that is
less adaptive. The comparison below runs three decay rates so the two failure
directions and the working case can be seen against each other.

### Choosing a target bar size

There is no optimal target. Choose it against:
- Desired bar frequency (more bars = better normality, but more noise)
- Trading horizon (intraday needs more bars than swing)
- Signal strength vs statistical properties tradeoff

## 6. Comparing Three Approaches: α-Based, Fixed, and Window-Based

The ml4t-engineer library provides three different implementations:

1. **α-Based (AFML)**: Exponential decay for E[T] and P[b=1] - requires careful α tuning
2. **Fixed Threshold**: No adaptation - simplest and most predictable
3. **Window-Based**: Rolling window adaptation - bounded drift

Let's compare them on the same data.

```python
from ml4t.engineer.bars import (
    FixedTickImbalanceBarSampler,  # Fixed threshold
    TickImbalanceBarSampler,  # α-based
    WindowTickImbalanceBarSampler,  # Window-based
)

# Test parameters - match the handoff comparison
COMPARE_ET = 1000
COMPARE_THRESHOLD = 100  # For fixed


# Track how E[T] drifts for each method
def measure_et_drift(bars: pl.DataFrame) -> float:
    """Measure E[T] drift (last / first expected_t)."""
    if "expected_t" not in bars.columns or len(bars) == 0:
        return 1.0  # No drift for fixed or empty bars
    first = bars["expected_t"][0]
    last = bars["expected_t"][-1]
    return last / first if first > 0 else 1.0
```

```python
print("=" * 70)
print("COMPARING THREE TICK IMBALANCE BAR APPROACHES")
print("=" * 70)

comparison = []

# 1. α-based with different alphas
for alpha in [0.001, 0.01, 0.1]:
    bars = TickImbalanceBarSampler(
        expected_ticks_per_bar=COMPARE_ET,
        alpha=alpha,
        min_bars_warmup=100,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"α={alpha}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"α-based α={alpha}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )
```

```python
# 2. Fixed threshold
for thresh in [50, 100, 200]:
    bars = FixedTickImbalanceBarSampler(threshold=thresh).sample(trades)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"fixed={thresh}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": 1.0,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Fixed thresh={thresh}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift=N/A"
    )
```

```python
# 3. Window-based with different tick windows
for tick_win in [2000, 5000, 10000]:
    bars = WindowTickImbalanceBarSampler(
        initial_expected_t=COMPARE_ET,
        bar_window=10,
        tick_window=tick_win,
    ).sample(trades)
    drift = measure_et_drift(bars)
    bar_stats = compute_stats(bars)
    comparison.append(
        {
            "method": f"window={tick_win}",
            "n_bars": len(bars),
            "avg_ticks": len(trades) / len(bars) if len(bars) > 0 else 0,
            "et_drift": drift,
            "jb": bar_stats["jarque_bera"],
            "ac1": bar_stats["autocorr_1"],
        }
    )
    print(
        f"Window tick_win={tick_win}: {len(bars):>4} bars, "
        f"avg_ticks={len(trades) / len(bars) if len(bars) > 0 else 0:>7.0f}, E[T] drift={drift:.2f}x"
    )
```

```python
# Summary table
compare_df = pl.DataFrame(comparison)
print("\n" + "=" * 70)
print("COMPARISON SUMMARY")
print("=" * 70)
compare_df = compare_df.with_columns(
    [
        pl.col("avg_ticks").round(0).cast(pl.Int64),
        pl.col("et_drift").round(2),
        pl.col("jb").round(1),
        pl.col("ac1").round(3),
    ]
)
print(compare_df)
```

## 7. Recommendations

| Use case | Sampler | Why |
|----------|---------|-----|
| Production | `FixedTickImbalanceBarSampler` | The threshold cannot drift, so bar size is predictable |
| Research | `TickImbalanceBarSampler` with a slow decay | Follows the textbook scheme while damping the feedback loop |

The table above carries two different things and they answer two different questions.

The bar count and the average bar size say what the sampler *produced*. Both depend on
the trade-sign sequence as well as on the threshold, so a bar shorter than the target
is not on its own evidence that the threshold moved - a run of one-sided signs fills
any threshold quickly, and the warm-up bars are short before adaptation has begun.

`et_drift` gets closer, and it is worth being exact about what it is: the sampler's
expected bar size at the last bar over its value at the first. That is one of the two
factors the threshold is built from - the threshold is the expected bar size times the
expected imbalance, and both adapt - so `et_drift` far from one is evidence the
adaptation moved, while `et_drift` near one is not evidence that it did not: the two
factors can move against each other, and an endpoint ratio says nothing about what
happened in between.

The threshold itself is what settles it, and the sampler carries the pieces. Reading
its recorded expected bar size and expected imbalance bar by bar, rather than at the
endpoints, is what distinguishes a threshold that grew, one that fell away, and one
that oscillated to a similar-looking endpoint.

None of this applies to the fixed sampler, which has no adapting threshold. What to
check there is whether the bar count it produced is the one its threshold was
calibrated for.

```python
print("\n" + "=" * 70)
print("NOTEBOOK SUMMARY")
print("=" * 70)
print("\nTIBs:")
print(tib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nVIBs:")
print(vib_df.select(["expected_t", "n_bars", "jarque_bera", "autocorr_1", "variance_ratio_5"]))
print("\nThree-Approach Comparison:")
print(compare_df)
print("\nNotebook completed.")
```
![notebook output](figures/p1_1.png)

출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT

이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.