Перейти к содержимому
Все документы библиотеки

Издержки бэктеста акций: фильтры и частота ребалансировки

Код Machine Learning for Trading

Сводка

В ноутбуке изучается влияние транзакционных издержек на часто ребалансируемую стратегию NASDAQ-100. Сначала к существующим бэктестам до учёта издержек применяются предположения о расходах в базисных пунктах, чтобы показать, как меняется коэффициент Шарпа при росте издержек. Затем сравниваются версии общих равновзвешенных стратегий top-k для полного набора активов и после фильтрации по издержкам, чтобы выделить эффект выбора более дешёвых в торговле инструментов. Отдельный перебор меняет частоту ребалансировки и издержки на акцию, исследуя влияние более редкой торговли на результаты и число сделок.

В анализе выбор набора активов и частота ребалансировки рассматриваются как отдельные способы уменьшить издержки: первый меняет торгуемые ценные бумаги, второй — частоту оборота позиций. Особое внимание уделяется безубыточному уровню издержек как полезному требованию к исполнению. В переборе частоты издержки упрощённо делятся на комиссию и проскальзывание по сумме на акцию; это отличается от производственной модели с фиксированной комиссией и половиной спреда для каждого актива. Поэтому частоты сравниваются при общем упрощённом предположении, а не на основе фактических издержек исполнения.

Ключевые идеи

  • Транзакционные издержки могут преобладать в стратегии, если ожидаемая доходность на одну ребалансировку невелика.
  • Фильтр торгуемости меняет и издержки на сделку, и состав бумаг, определяющих оборот портфеля.
  • Снижение частоты ребалансировки распределяет издержки входа на более длительный период удержания.
  • Изменение одного фактора издержек за раз помогает определить, чем обусловлено улучшение результатов.
  • Перебор издержек на акцию — упрощённое сравнение, которое не воспроизводит производственную модель издержек.

Теги

Полный текст
# 17_costs.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # NASDAQ-100 Microstructure: Costs
#
# **Chapter 18 — Transaction Costs and Execution**
#
# This is the primary cost-analysis notebook for the NASDAQ-100 case study.
#
# Trading costs enter a strategy once per trade, so what they take out over a
# year depends on how often it trades and on how expensive each name is to trade.
# At a 15-minute rebalancing interval, a strategy crosses the spread many times a
# day on every position it holds, and the expected return over each of those
# intervals is a fraction of a basis point. The two quantities are close enough
# in size that the cost assumption decides the outcome.
#
# The notebook measures two levers against that, separately, so each can be
# attributed:
#
# 1. **Which names to trade.** Restricting to the cheapest-to-trade names
#    changes both what is paid per trade and how much trading the ordering
#    provokes, because the expensive names are also the ones whose ranking moves
#    around most.
# 2. **How often to trade.** Rebalancing less frequently spreads each entry cost
#    over a longer holding period. The signal is unchanged; what changes is how
#    many times it is acted on.
#
# The notebook has three parts:
# - **Sections 1-3**: Standard bps cost grid on the leading pre-cost runs, which
#   on this case study are the screened baseline-stage arms (Section 1 says why the
#   allocation and risk-overlay stages contribute none), tracing the
#   Sharpe-vs-cost decay curve.
# - **Section 4**: Full universe vs the cost-feasible screen — the first lever,
#   read off existing registry rows for the equal-weight top-k arms, which are
#   the arms both universes carry.
# - **Section 5**: Cadence × per-share cost sweep — the second lever and the
#   publication finding. Uses a per-share cost model ($/share, not bps), more
#   realistic for equities, swept across rebalance frequencies.
#
# **Learning Objectives:**
# 1. Sweep a cost grid over existing backtests to locate the cost level at
#    which a strategy stops being profitable
# 2. Separate the effect of restricting the universe from the effect of
#    rebalancing less often, by varying one at a time
# 3. Read a two-dimensional sweep of rebalancing frequency against per-share
#    cost, and say what execution quality each region of it assumes
#
# **Book Reference:** Chapter 18, Sections 18.2–18.5
#
# **Prerequisites:** [`16_risk_management`](16_risk_management.ipynb), and through it
# [`15_portfolio_management`](15_portfolio_management.ipynb) and [`14_backtest`](14_backtest.ipynb).
# This is the last stage that selects, so it runs after all three and draws from all of
# them.

# %%
"""NASDAQ-100 Microstructure: Costs."""

import json
import sqlite3
import time

import polars as pl

from case_studies.research import open_study
from case_studies.utils.backtest_loaders import (
    get_backtest_config,
    load_backtest_prices_for,
    warmup_periods_for,
)
from case_studies.utils.backtest_presets import (
    build_backtest_spec,
    clone_backtest_spec,
    ensure_backtest_spec,
    prediction_age_declaration,
    set_backtest_costs_bps,
    strategy_view,
    traded_universe_declaration,
)
from case_studies.utils.backtest_runner import run_backtest
from case_studies.utils.notebook_contracts import excluded_families
from case_studies.utils.registry import read_predictions, resolve_best_backtest_runs
from case_studies.utils.sweep_config import (
    get_cadence_sweep,
    get_cost_grid_bps,
    get_cost_grid_half_spread_usd,
    get_top_n_predictions,
    get_universe_filters_for,
)
from case_studies.utils.uncertainty import STAGE_SEQUENCE
from utils.paths import get_case_study_dir
from utils.style import show_with_alt

# %% tags=["parameters"]
CASE_STUDY_ID = "nasdaq100_microstructure"
LABEL = ""
MAX_SYMBOLS = 0
TOP_N_COMBOS = None
# Both names stay bound here although nothing below reads them: that is what makes the harness
# force preview and supply a workspace - `_declares_tier_and_workspace` in `tests/pm_helpers.py`
# looks for exactly this pair. Without them the canonical branch regenerates in place, which
# needs generated-artifact symlinks a CI checkout does not have.
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""

# %%
# A reduced run is a preview run. Refused on the canonical tier so a narrowed result can
# never land in the registry the book's numbers come from, and so the two can never sit in
# one registry to be ranked against each other: `resolve_best_backtest_runs` takes the top
# Sharpe over every backtest at a stage, and a Sharpe earned over a handful of names would
# outrank one earned over the whole panel. `us_equities_panel` 16 through 19 already refuse
# the parameter this way, and `canonically_refused_parameters` reads the refusal out of the
# source, so the canonical fixture path drops the name rather than handing the notebook
# something its first cell raises on (ml4t/agent-workspace#911).
if EXECUTION_TIER == "canonical" and MAX_SYMBOLS:
    raise ValueError(
        "MAX_SYMBOLS narrows the universe this run trades, which makes it a different "
        "portfolio from the declared one and gives it its own backtest identity "
        "(ml4t/agent-workspace#911). A canonical run trades the declared universe: set "
        "MAX_SYMBOLS=0, or run under EXECUTION_TIER='preview' with a WORKSPACE."
    )

# %% [markdown]
# The study is opened before anything resolves a path or reads the registry. Opening it
# activates a root and rewrites `ML4T_OUTPUT_DIR` process-wide, and every later
# `get_case_study_dir`, prediction read and registry write resolves against that variable. A
# `CASE_DIR` bound before this line points at the released registry while this notebook writes
# to the workspace, and the two never meet: the sweep finds nothing registered and every reader
# scoped to hashes from the other root comes back empty.

# %%
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)

CASE_DIR = get_case_study_dir(CASE_STUDY_ID)
bt_config = get_backtest_config(CASE_STUDY_ID)
if TOP_N_COMBOS is None:
    TOP_N_COMBOS = get_top_n_predictions(CASE_STUDY_ID, "cost_sensitivity")
if not LABEL:
    LABEL = bt_config.primary_label

print(f"Case study: {CASE_STUDY_ID}, label: {LABEL}")

COST_GRID_BPS = get_cost_grid_bps(CASE_STUDY_ID)

if excluded_families(CASE_STUDY_ID):
    print(
        "Active-model filter: excluding "
        f"{', '.join(sorted(excluded_families(CASE_STUDY_ID)))} pending corrected reruns"
    )

# %% [markdown]
# ## 1. Load the leading pre-cost runs
#
# We load the highest-Sharpe validation backtests on the universe this case study treats as
# canonical, and trace how their Sharpe-vs-cost curve behaves before the two recovery
# mechanisms - the screen in Section 4 and the cadence in Section 5 - are applied.
#
# **The pool is every stage a selected configuration can come from, not just `allocation`.** A risk
# overlay is a strategy in its own right: `16_risk_management` registers it at
# `stage='risk_overlay'` with its own Sharpe, and it is a candidate to carry the case study.
# Pricing only the allocation rows would put a cost curve in the chapter for a strategy the case
# study does not select whenever an overlay outranks its own parent, which is the ordinary case -
# four of the seven completed case studies have a `risk_overlay` as their rank-1 validation
# configuration. The stages come from `STAGE_SEQUENCE` rather than a tuple typed here, so the pool
# cannot drift from the library when a stage is added.
#
# `cost_sensitivity` is the one member excluded, because that is the stage this notebook
# writes: including it would re-price rows that already carry a cost model.
#
# **The pool is also pinned to the canonical universe.** `setup.yaml` declares
# `universe_filter: cost_feasible` and says in the same breath that the full-universe variant
# "is NOT a canonical rank-1 / cohort / DSR candidate". Widening the stages without pinning the
# universe would admit exactly that variant, because the baseline stage holds both screened and
# full-universe runs and the full-universe ones are not screened out anywhere else - the cost
# curve would then price a strategy the case study excludes by declaration. The filter is read
# out of each run's own `spec_json`, the same place `derived_tables_off_canonical_universe`
# reads it, so the pool cannot disagree with the sweep that produced the runs.
#
# Section 4 is unaffected and has to be: the full-versus-screened contrast is the one place
# the excluded variant belongs, and it reads `universe_filter` in its own query rather than
# through this pool.
#
# This is also why the notebook is numbered after `16_risk_management` rather than before it.
# Run the other way round, the overlay rows do not exist yet and the pool is `allocation`
# whatever it declares.
#
# **On this case study the two arguments above point in opposite directions, and the pin
# wins.** `15_portfolio_management` builds its allocation specs with no `universe_filter` and
# `16_risk_management` clones those specs, so every allocation and risk-overlay row this
# registry holds is a full-universe row and the pin drops all of them. The pool that reaches
# the cost grid is the baseline stage's screened rows. That is the declared design rather than a
# gap - 15 exists to show that allocation does not rescue the every-bar full-universe
# strategy, and a negative result is not a candidate to carry the case study - but it does mean
# the risk-overlay argument above is about the other eight case studies and not about this one.

# %%
PRE_COST_STAGES = tuple(stage for stage in STAGE_SEQUENCE if stage != "cost_sensitivity")
CANONICAL_UNIVERSE = get_universe_filters_for(CASE_STUDY_ID)[0]


def _on_canonical_universe(frame: pl.DataFrame, stage: str = "upstream") -> pl.DataFrame:
    """Drop runs selected under a universe this case study does not treat as canonical.

    `None` means the case study pins no universe, and then every run qualifies - the filter
    has to be a no-op there rather than dropping everything, because a spec that predates the
    universe axis carries no filter at all.
    """
    if CANONICAL_UNIVERSE is None:
        return frame
    # An empty resolver result carries no columns at all, so reading `spec_json` off it raised
    # `ColumnNotFoundError: "spec_json" not found`, naming a column rather than the absence that
    # produced it. Returning the empty frame is NOT the fix: it made this notebook exit 0 having
    # registered nothing, which is the same absence wearing a success. Measured 2026-09-09 on the
    # smoke chain - 202 signal backtests registered, none carrying `universe_filter`, and this
    # notebook reported no error at all.
    if frame.is_empty():
        raise RuntimeError(
            f"no {stage} backtests are registered for {CASE_STUDY_ID}, so there is nothing "
            "to price. Run 14_backtest through 16_risk_management against this registry first."
        )
    keep = [
        strategy_view(json.loads(spec)).get("signal", {}).get("universe_filter")
        == CANONICAL_UNIVERSE
        for spec in frame["spec_json"]
    ]
    return frame.filter(pl.Series(keep, dtype=pl.Boolean))


def resolve_pre_cost_runs(top_n: int) -> pl.DataFrame:
    """The highest-Sharpe validation runs across every stage a selected configuration may come from.

    Each stage is asked for its whole ranked list and the pool is sorted afterwards, rather
    than taking `top_n` from each and merging them: truncating first lets one stage's leader
    hold a slot that a better run in another stage should have had, and at `top_n=1` that
    drops a whole stage from consideration instead of falling through to the next candidate.
    """
    ranked = [
        frame.with_columns(pl.lit(stage).alias("pool_stage"))
        for stage, frame in (
            (
                stage,
                _on_canonical_universe(
                    resolve_best_backtest_runs(
                        CASE_STUDY_ID, LABEL, split="validation", stage=stage, top_n=1_000_000
                    ),
                    stage,
                ),
            )
            for stage in PRE_COST_STAGES
        )
        if not frame.is_empty()
    ]
    if not ranked:
        return pl.DataFrame()
    return (
        pl.concat(ranked)
        .sort("sharpe", descending=True)
        .unique("backtest_hash", maintain_order=True)
        .head(top_n)
    )


top_combos = resolve_pre_cost_runs(TOP_N_COMBOS)

if top_combos.is_empty():
    # Refused rather than reported. Every cell below is a no-op on an empty pool, so the
    # notebook would run to the end, register nothing and exit 0 - a chapter with no cost
    # curve in it and no error anywhere saying why.
    msg = (
        f"No {CANONICAL_UNIVERSE or 'full'}-universe validation backtests at any of "
        f"{', '.join(PRE_COST_STAGES)} for {CASE_STUDY_ID}/{LABEL}. The stages hold rows or "
        "`_on_canonical_universe` would have raised first, so what is missing is the "
        f"universe: nothing registered declares `signal.universe_filter = "
        f"{CANONICAL_UNIVERSE!r}`. Re-run 14_backtest, whose baseline pass is what registers "
        "them."
    )
    raise RuntimeError(msg)
for row in top_combos.iter_rows(named=True):
    spec = ensure_backtest_spec(
        CASE_STUDY_ID,
        bt_config,
        json.loads(row["spec_json"]),
        prices=load_backtest_prices_for(
            CASE_STUDY_ID,
            LABEL,
            split="validation",
            warmup_periods=warmup_periods_for(CASE_STUDY_ID),
            max_symbols=MAX_SYMBOLS,
        ),
        prediction_hash=row["prediction_hash"],
        initial_cash=bt_config.initial_cash,
    )
    alloc = strategy_view(spec).get("allocation", {}).get("method", "equal_weight")
    # The stage is printed because it is the thing that changed: a `risk_overlay` configuration
    # and its `allocation` parent share a prediction hash, so nothing else in this line
    # distinguishes the overlaid run from the un-overlaid one it was built on.
    print(
        f"  Sharpe={row['sharpe']:.3f}  stage={row['pool_stage']}  alloc={alloc}  "
        f"bt_hash={row['backtest_hash'][:8]}"
    )

# %%
prices = load_backtest_prices_for(
    CASE_STUDY_ID,
    LABEL,
    split="validation",
    warmup_periods=warmup_periods_for(CASE_STUDY_ID),
    max_symbols=MAX_SYMBOLS,
)
print(f"Prices: {len(prices):,} rows, {prices['symbol'].n_unique()} assets")

# %% [markdown]
# ## 2. Cost Grid Sweep
#
# For each top combo, re-run the backtest at different total cost levels
# (commission + slippage combined). The grid spans from near-zero to levels
# that exceed the signal entirely, tracing the full decay curve.
#
# At 15-minute cadence the strategy is offered 20 decision slots a trading day
# (Section 5 measures the grid at every cadence), so even 1 bps per leg compounds
# to a large annual drag. Where the breakeven actually falls is what the curve
# below reports; the reason to expect it low is arithmetic rather than a result,
# and the two mechanisms that move it are the ones the rest of this chapter
# measures - trading fewer names, and trading them less often.

# %%
n_total = len(top_combos) * len(COST_GRID_BPS) if not top_combos.is_empty() else 0
n_done = 0
t0 = time.time()
# The rows this run registers, so section 3 can plot its own curve rather than everything the
# `cost_sensitivity` stage has ever held. Collected here because the hash is only known after
# the run returns.
swept_hashes: list[str] = []
# Grid points that did not produce a backtest. A cost curve with holes is not a slower
# version of the curve - it is a different one, and the breakeven read off it states a cost
# level the strategy was never tested at. Collected rather than only printed, so the
# rendering below can refuse instead of interpolating across the gap.
failed_points: list[str] = []

for combo_row in top_combos.iter_rows(named=True):
    pred_hash = combo_row["prediction_hash"]
    base_spec = ensure_backtest_spec(
        CASE_STUDY_ID,
        bt_config,
        json.loads(combo_row["spec_json"]),
        prices=prices,
        prediction_hash=pred_hash,
        initial_cash=bt_config.initial_cash,
    )
    alloc_method = strategy_view(base_spec).get("allocation", {}).get("method", "equal_weight")
    predictions = read_predictions(CASE_STUDY_ID, pred_hash)

    for cost_bps in COST_GRID_BPS:
        n_done += 1
        spec = set_backtest_costs_bps(
            clone_backtest_spec(base_spec),
            commission_bps=cost_bps / 2,
            slippage_bps=cost_bps / 2,
        )
        spec["chapter"] = "ch18"
        try:
            result = run_backtest(
                CASE_STUDY_ID,
                pred_hash,
                spec,
                prices=prices,
                predictions=predictions,
                label=LABEL,
                register=True,
                initial_cash=bt_config.initial_cash,
                calendar=bt_config.calendar,
            )
            if result.backtest_hash:
                swept_hashes.append(result.backtest_hash)
            if cost_bps % 10 == 0:
                print(
                    f"  [{n_done}/{n_total}] {alloc_method} @ {cost_bps}bps: "
                    f"Sharpe={result.metrics.get('sharpe', 0):.3f}"
                )
        except Exception as e:
            failed_points.append(f"{alloc_method} @ {cost_bps}bps: {e}")
            print(f"  [{n_done}/{n_total}] {alloc_method} @ {cost_bps}bps: FAILED — {e}")

# %%
elapsed = time.time() - t0
print(f"Cost sweep complete: {n_done} backtests in {elapsed:.0f}s")
if failed_points:
    raise RuntimeError(
        f"{len(failed_points)} of {n_total} cost-grid points did not produce a backtest, so "
        "the decay curve below would be drawn through the gaps and the breakeven read off it "
        "would name a cost level nothing was tested at. Fix the failures and re-run; the "
        "points that succeeded are registered and will be reused.\n  " + "\n  ".join(failed_points)
    )

# %% [markdown]
# ## 3. Cost Sensitivity Analysis
#
# This section is **read-only** — queries the registry for cost-sensitivity
# results and computes breakeven levels.
#
# The number to take from the curve is the **breakeven** - the cost level at which
# Sharpe crosses zero - because it states the execution quality the strategy requires
# rather than the profit it made under one assumption. It is interpolated below from the
# grid points and printed. No band is predicted here: a prediction written beside the
# calculation is read as its confirmation.
#
# Two features of the shape are worth reading beside it. How steeply the curve falls from
# the zero-cost end says how much of the gross result was ever available - a strategy whose
# Sharpe halves by 1 bps was never trading on much. And a flat stretch, if there is one, is
# the range over which execution quality does not change the answer, the only part of the
# axis where a cost assumption can be wrong without mattering.
#
# The grid is `backtest.sweep.cost_grid_bps`, 0 to 50 bps per leg. What it is measured
# against is not a bps figure: every baseline-stage row in this registry carries
# `commission.model = "per_share"` at $0.0035 a share, and the `rate` fields beside it
# belong to the percentage model that is not the one in force. `cost_sensitivity` returns
# only the rows this notebook re-ran under the percentage model, which is what puts them
# on one axis at all.

# %%
from case_studies.utils.backtest_explorer import BacktestExplorer

explorer = BacktestExplorer(CASE_STUDY_ID)

# %% [markdown]
# **The curve is scoped to the rows this run just registered.** `cost_sensitivity()` unscoped
# returns every row the stage has ever held - configurations selected earlier, superseded
# generations, and
# the full-universe rows section 4 registers on purpose. Plotting those together produces one
# line per allocator drawn through several strategies at once, which is not a Sharpe-versus-cost
# curve for anything. `backtest_explorer.cost_sensitivity`'s own docstring names this case study
# as one that must scope, and it was not scoping.

# %%
if not swept_hashes:
    # Falling back to an unscoped read here would be the worst of both: the curve reappears,
    # mixed across every generation the stage has ever held, precisely on the runs where the
    # sweep above registered nothing and there is no curve to draw. An empty frame says the
    # sweep produced nothing, which is what happened.
    print("The sweep above registered no cost rows, so there is no curve for this run to plot.")
    cost_df = pl.DataFrame()
else:
    cost_df = explorer.cost_sensitivity(backtest_hashes=swept_hashes)

if not cost_df.is_empty():
    import matplotlib.pyplot as plt

    fig, ax = plt.subplots(figsize=(10, 6))

    for alloc in cost_df["allocator"].unique().sort().to_list():
        subset = cost_df.filter(pl.col("allocator") == alloc).sort("cost_bps")
        ax.plot(subset["cost_bps"].to_list(), subset["sharpe"].to_list(), marker="o", label=alloc)

    ax.axhline(0, color="gray", linestyle="--", alpha=0.5)
    ax.axhline(0.5, color="gray", linestyle=":", alpha=0.5)
    ax.set_xlabel("Total Cost (bps per leg)")
    ax.set_ylabel("Net Sharpe Ratio")
    ax.set_title("Sharpe Decay Under Transaction Costs")
    ax.legend()
    # `utils/style` and `matplotlibrc` both set `figure.constrained_layout.use`, so
    # `tight_layout()` warns and fights the layout engine already running; `fig.show()` warns
    # that a non-interactive canvas cannot be shown and publishes no alt text. Measured: two
    # UserWarnings per figure. The six notebooks of this case study already at `done` use
    # `show_with_alt` and neither of the other two calls.
    show_with_alt(
        fig,
        "Line chart of net Sharpe against transaction cost. The horizontal axis is total "
        "cost in basis points per leg and the vertical axis is net Sharpe ratio. One line "
        "with circular markers per allocator, named in the legend, tracing that allocator's "
        "Sharpe as cost rises. Two horizontal reference lines mark zero Sharpe (dashed) and "
        "0.5 (dotted), so the cost at which a line crosses each can be read off.",
    )
else:
    print("No cost sensitivity data in registry")

# %%
# The breakeven, interpolated between the grid points that straddle zero. The grid is
# coarse (0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50), so the crossing is reported as an
# interval as well as a point: a reader who needs it tighter needs more grid, not more
# precision in this arithmetic.
if not cost_df.is_empty():
    # `cost_sensitivity` returns cost_bps, sharpe, max_drawdown and allocator, and no
    # identity, so a curve here is only a curve while one configuration contributes each
    # (allocator, cost_bps) point. `top_n_predictions.cost_sensitivity` is 1 on this case
    # study, so that holds; raise it above 1 and two selected configurations sharing an
    # allocator pool into one line, and an interpolation across them describes neither.
    # Checked rather than assumed, because the collision is invisible in the plot.
    _dupes = (
        cost_df.group_by(["allocator", "cost_bps"])
        .agg(pl.len().alias("n"))
        .filter(pl.col("n") > 1)
        .sort("n", descending=True)
    )
    if not _dupes.is_empty():
        msg = (
            f"{_dupes.height} (allocator, cost) point(s) carry more than one backtest, so "
            "these curves pool configurations that share an allocator and no breakeven "
            f"read off them belongs to a single strategy. Worst: {_dupes.row(0)}. Scope "
            "the cost sweep to one configuration, or extend `cost_sensitivity` to return "
            "the identity so the curves can be split by it."
        )
        raise RuntimeError(msg)

    _breakeven_rows = []
    for _alloc in cost_df["allocator"].unique().sort().to_list():
        _curve = cost_df.filter(pl.col("allocator") == _alloc).sort("cost_bps")
        _x = _curve["cost_bps"].to_list()
        _y = _curve["sharpe"].to_list()
        if not _y or _y[0] <= 0:
            _breakeven_rows.append(
                {
                    "allocator": _alloc,
                    "breakeven_bps": "at or below the lowest grid point",
                    "bracket": f"<= {_x[0]:g}" if _x else "n/a",
                }
            )
            continue
        _cross = next((i for i in range(1, len(_y)) if _y[i] <= 0), None)
        if _cross is None:
            _breakeven_rows.append(
                {"allocator": _alloc, "breakeven_bps": "above the grid", "bracket": f"> {_x[-1]:g}"}
            )
            continue
        _x0, _x1, _y0, _y1 = _x[_cross - 1], _x[_cross], _y[_cross - 1], _y[_cross]
        _be = _x0 + (_x1 - _x0) * (_y0 / (_y0 - _y1)) if _y0 != _y1 else _x1
        _breakeven_rows.append(
            {"allocator": _alloc, "breakeven_bps": f"{_be:.2f}", "bracket": f"{_x0:g}-{_x1:g}"}
        )
    print("Breakeven cost per leg (Sharpe crosses zero):")
    print(pl.DataFrame(_breakeven_rows))

# %% [markdown]
# ## 4. Full Universe vs the Cost-Feasible Screen
#
# The first lever is which names to trade. Half the bid-ask spread is what a
# trade pays to cross it, and that varies widely across the panel: the same
# strategy pays several times more per trade in the least liquid names than in
# the most liquid ones. The cost-feasible universe keeps the cheapest-to-trade
# names, fixed per split so the screen cannot use information from the window it
# is evaluated on.
#
# This section reads the equal-weight top-k arms on both universes directly from
# the registry, running no new backtests, and compares the outcome alongside the
# trade count. Equal weight and not the featured slot design, because the slot
# design is registered on the screened universe alone and a screen effect needs
# both sides; which arms exist where is set by `baseline_schemes` and
# `reference_schemes` in `config/setup.yaml`.
#
# Reporting outcome and trade count together matters: the screen changes which
# names can be held, so it changes how much trading the ordering provokes as well
# as what each trade costs, and the two effects are not separable from the outcome
# alone.

# %%
# Read on the equal-weight arms, not the slot design, and matched pair by pair.
#
# Which arms exist on which universe is decided by `config/setup.yaml`, not here:
# `baseline_schemes` runs ew_top5/10/20 on `baseline_universe: cost_feasible` for every
# pass-1 prediction, and `reference_schemes` runs the same three on
# `reference_universe: full` for the pass-2 predictions only. The slot design this
# section used to query is registered on the screened universe alone, so that query
# returned one row and the derived line below it was skipped, leaving a one-sided table
# under prose describing two sides.
#
# Matched arm by arm, because the two universes do not cover the same predictions: the
# screened side carries every pass-1 prediction and the full side only the pass-2
# survivors. Averaging each side over its own population would compare 510 backtests
# against 10 and attribute the difference to the screen, when most of it is the
# difference between the two populations.
#
# The match key is the whole signal specification with `universe_filter` removed, not
# `(prediction_hash, top_k)`. Those two fields do not identify an arm: pass 2 registers
# `equal_weight_top_k` on the screened universe at both `long_short: true` and a
# `long_only` variant, the full-universe reference carries the long-short one only, and
# a join on prediction and concentration alone therefore pairs one full row with two
# screened rows. The difference it reports would then mix the universe screen with a
# change of trading direction. Keyed on the sorted items of the signal dict so the match
# does not depend on the order SQLite happens to serialize the object in.
#
# Restricted further to arms carrying no allocator, because the signal stage holds two
# kinds of row and only one of them belongs here. `15_portfolio_management` prices
# `equal_weight` as one of its three allocators, and an equal-weight allocation of a
# top-k basket is the same selection as the signal arm, so those cells register at
# `stage: signal` rather than writing an allocation row. They are not duplicates of the
# ch16 rows: measured 2026-09-14 across all 18 colliding pairs, the two sides carry
# identical `num_trades` and different Sharpes - ew_top5 on the full universe is -4.007
# from ch16 against -4.056 from ch17, ew_top10 -8.821 against -8.893, ew_top20 -12.667
# against -12.754. Same entries and exits, different position weights. So the allocated
# row is a real and distinct backtest, and the fix is to leave it to the allocation
# stage that produced it rather than to merge the pair or pick one.
#
# Without the restriction the join below is not one to one and the assertion that
# follows raises, which is how this was found.
conn = sqlite3.connect(str(CASE_DIR / "run_log" / "registry.db"))
_arms = pl.read_database(
    """
    SELECT
        COALESCE(json_extract(br.spec_json, '$.strategy.signal.universe_filter'),
                 'full')                                                  AS universe,
        br.prediction_hash                                                AS prediction_hash,
        tr.family                                                         AS family,
        json_extract(br.spec_json, '$.strategy.signal')                   AS signal_json,
        bm.sharpe                                                         AS sharpe,
        bm.num_trades                                                     AS num_trades
    FROM backtest_runs br
    JOIN backtest_metrics bm ON br.backtest_hash = bm.backtest_hash
    JOIN prediction_sets ps ON br.prediction_hash = ps.prediction_hash
    JOIN training_runs tr ON tr.training_hash = ps.training_hash
    WHERE br.stage = 'signal' AND ps.split = 'validation'
      AND json_extract(br.spec_json, '$.strategy.signal.method') = 'equal_weight_top_k'
      AND json_extract(br.spec_json, '$.strategy.allocation.method') IS NULL
      AND bm.sharpe IS NOT NULL
    """,
    connection=conn,
    schema_overrides={"sharpe": pl.Float64, "num_trades": pl.Float64},
)
conn.close()


def _arm_key(signal_json: str) -> str:
    """The arm a backtest ran, with the universe taken out of it."""
    signal = json.loads(signal_json)
    signal.pop("universe_filter", None)
    return json.dumps(signal, sort_keys=True)


_arms = _arms.with_columns(
    pl.col("signal_json").map_elements(_arm_key, return_dtype=pl.String).alias("arm"),
)


def _pairs_on_both(frame: pl.DataFrame) -> pl.DataFrame:
    """The (prediction, arm) pairs this frame carries on the full universe and on the screen."""
    return (
        frame.filter(pl.col("universe") == "full")
        .select("prediction_hash", "arm")
        .join(
            frame.filter(pl.col("universe") == "cost_feasible").select("prediction_hash", "arm"),
            on=["prediction_hash", "arm"],
            how="inner",
        )
    )


# The section's claim is about the screen and not about a model family, so the family is a
# choice about which comparison gets published rather than a condition of the comparison
# being valid. gbm is asked for first because that is what previous renders carried and a
# family that changes between renders changes the number under the prose. It used to be a
# `tr.family = 'gbm'` clause in the query above, which is the same choice written as though
# it were a requirement: a run that fits no gbm at all - the pull-request fixture runs
# 06_linear and 07_gbm on the weekly tier and skips them - then matched nothing, and the
# refusal below reported a sweep that had not reached pass 2 when what had happened is that
# the one family it would look at was absent. Any family carrying the same (prediction, arm)
# on both universes answers the same question.
_candidate_families = ["gbm"] + sorted(set(_arms["family"].unique()) - {"gbm"})
_family = next(
    (
        f
        for f in _candidate_families
        if not _pairs_on_both(_arms.filter(pl.col("family") == f)).is_empty()
    ),
    "gbm",
)
_arms = _arms.filter(pl.col("family") == _family)
print(f"Section 4 compares the {_family} family across the two universes")

_full = _arms.filter(pl.col("universe") == "full").select(
    "prediction_hash", "arm", "sharpe", "num_trades"
)
_screened = _arms.filter(pl.col("universe") == "cost_feasible").select(
    "prediction_hash", "arm", "sharpe", "num_trades"
)
# One backtest per (prediction, arm, universe) is what the registry's identity
# guarantees; asserted rather than assumed, because a duplicate would silently weight
# one prediction twice in the averages below.
for _side_name, _side in (("full", _full), ("cost_feasible", _screened)):
    _dupes = _side.group_by("prediction_hash", "arm").len().filter(pl.col("len") > 1)
    if not _dupes.is_empty():
        msg = (
            f"{_dupes.height} (prediction, arm) pairs appear more than once on the "
            f"{_side_name} universe, so the match would not be one to one"
        )
        raise RuntimeError(msg)

_matched = _full.join(_screened, on=["prediction_hash", "arm"], how="inner", suffix="_screened")
screen_compare = (
    pl.concat(
        [
            _matched.select(
                pl.lit("full").alias("universe"),
                pl.col("sharpe"),
                pl.col("num_trades"),
            ),
            _matched.select(
                pl.lit("cost_feasible").alias("universe"),
                pl.col("sharpe_screened").alias("sharpe"),
                pl.col("num_trades_screened").alias("num_trades"),
            ),
        ]
    )
    .group_by("universe")
    .agg(
        n_arms=pl.len(),
        avg_sharpe=pl.col("sharpe").mean().round(3),
        min_sharpe=pl.col("sharpe").min().round(3),
        max_sharpe=pl.col("sharpe").max().round(3),
        avg_trades=pl.col("num_trades").mean().round(0),
    )
    .sort("universe", descending=True)
)
print(f"{_matched.height} (prediction, arm) pairs registered on both universes")
print(screen_compare)

# %%
# The section's claim is a difference between two universes, so one universe is not a
# weaker version of it and the derived line must not be skipped when it finds one. A
# one-row table under prose describing two sides is the failure this refusal exists to
# stop, and it is the state the previous query was in.
_universes = set(screen_compare["universe"].to_list())
if _universes != {"full", "cost_feasible"}:
    msg = (
        "section 4 compares one arm across two universes and the matched set offers "
        f"{sorted(_universes) or 'none'}. It reads stage='signal', split='validation', "
        f"family={_family!r} (chosen from {_candidate_families}), "
        "signal.method='equal_weight_top_k', and needs the same (prediction, arm) pair "
        "registered on both the full universe and cost_feasible. `baseline_schemes` and "
        "`reference_schemes` in config/setup.yaml decide that; they currently agree on "
        "ew_top5/10/20, so an empty match means the sweep did not reach pass 2."
    )
    raise RuntimeError(msg)

full_row = screen_compare.filter(pl.col("universe") == "full")
screened_row = screen_compare.filter(pl.col("universe") == "cost_feasible")
d_sharpe = screened_row["avg_sharpe"][0] - full_row["avg_sharpe"][0]
# Screened over full, so the direction reads off the number: above 1 the screen traded
# more. A full-over-screened ratio was the previous form and it reads as a reduction
# whichever way the trade count moved.
trade_ratio = screened_row["avg_trades"][0] / max(full_row["avg_trades"][0], 1)
print(
    f"Screen moves avg Sharpe by {d_sharpe:+.2f} "
    f"({full_row['avg_sharpe'][0]:+.2f} to {screened_row['avg_sharpe'][0]:+.2f}) "
    f"and multiplies trades by {trade_ratio:.2f} "
    f"({full_row['avg_trades'][0]:.0f} to {screened_row['avg_trades'][0]:.0f})."
)

# %% [markdown]
# ### Reading the Screen's Effect
#
# Same arm, same model family, same validation window, and the same predictions on
# both sides. The only difference between the two rows is the tradeable universe, so
# the gap between them is what the screen did and not what the model did.
#
# Read the Sharpe difference and the trade ratio together. The screen changes which
# names can be held, so it changes how much trading the ordering provokes as well as
# what each trade costs, and a Sharpe difference alone cannot separate the two. The
# printed line above states both, computed from the table rather than described here,
# because a described ordering goes stale against the next rebuild while a computed
# one cannot.

# %% [markdown]
# ## 5. Cadence × Per-Share Cost Analysis
#
# The bps sweep above fixes the rebalancing cadence at 15 minutes. The
# cost-to-edge ratio depends on how often the strategy trades as well as on what
# each trade costs, and the cadence axis is the one that moves the first term.
#
# `resolve_decision_schedule` puts each cadence on the clock over the window this
# label is scored on, 10:31 to 15:43. Measured on it: 62 decision slots a day at
# five minutes, 20 at fifteen, 10 at thirty and 5 at sixty. A 9:30-to-16:00
# session would give 78, 26, 13 and 7; the scored window is shorter at both ends
# because a feature lookback opens it and the fifteen-minute forward label closes
# it. Holding longer amortizes the fixed per-trade cost over a larger expected
# return per period.
#
# This section sweeps **cadence × per-share spread** — the central exhibit
# for this case study. We use a **per-share cost model** rather than bps,
# because for equities the execution cost is a dollar amount per share (half
# the bid-ask spread plus commission), not a percentage of notional:
#
# The cost axis is expressed in cents per share. For a stock priced around one
# hundred dollars, a cent per share is close to one basis point of notional, so
# the axis spans from direct-market-access execution at the low end, through the
# effective spread a large-cap name typically quotes, to retail-quality
# execution at the high end. Naming the axis in cents rather than basis points
# keeps it comparable across stocks at different prices.
#
# **What the cost axis is.** The grid value `cost_ps` is the total per-share
# round-trip cost, split evenly between the engine's commission and slippage
# settings so the whole cost moves along one axis. The signal and allocation
# notebooks instead set a fixed per-share commission and add a per-asset
# half-spread measured from the liquidity profile, which are separate quantities
# on separate scales. The single axis here is what makes the interaction between
# rebalancing frequency and cost readable in one chart; it is a comparison across
# cadences under a simplified cost shape rather than a reproduction of the cost
# model used to produce the registered results.

# %%
from case_studies.utils.backtest_runner import normalize_prediction_columns
from case_studies.utils.registry import read_predictions

# Top engine baseline-stage prediction by Sharpe
db_path = CASE_DIR / "run_log" / "registry.db"
conn = sqlite3.connect(str(db_path))
cur = conn.cursor()
# The universe predicate is the same statement section 1 makes about the selected configuration
# pool, and it has to be made again here: this query picks its own row. Without it the cadence
# exhibit - the publication finding of this notebook - is built on whichever signal row ranks
# highest, which is the full-universe variant `setup.yaml` excludes from canonical candidacy
# whenever it wins.
cur.execute(
    """
SELECT br.prediction_hash, tr.family, tr.config_name, bm.sharpe
FROM backtest_runs br
JOIN backtest_metrics bm ON br.backtest_hash = bm.backtest_hash
JOIN prediction_sets ps ON br.prediction_hash = ps.prediction_hash
JOIN training_runs tr ON ps.training_hash = tr.training_hash
WHERE br.stage = 'signal'
AND json_extract(br.spec_json, '$.strategy.rebalance.mode') = 'engine'
AND tr.family != 'deep_learning'
AND (? IS NULL OR json_extract(br.spec_json, '$.strategy.signal.universe_filter') = ?)
ORDER BY bm.sharpe DESC
LIMIT 1
""",
    (CANONICAL_UNIVERSE, CANONICAL_UNIVERSE),
)
_row = cur.fetchone()
conn.close()

if _row is None:
    print(
        f"No baseline-stage engine backtest on the {CANONICAL_UNIVERSE or 'full'} universe. "
        "Skipping cadence sweep."
    )
    best_pred_hash = None
else:
    best_pred_hash = _row[0]
    best_source = f"{_row[1]}/{_row[2]}"
    print(f"Cadence sweep prediction: {best_source} (engine Sharpe={_row[3]:.3f})")

if best_pred_hash is not None:
    predictions_raw = normalize_prediction_columns(read_predictions(CASE_STUDY_ID, best_pred_hash))

    # Thin minute-level predictions to 15m for default cadence
    predictions_15m = predictions_raw.filter(
        (pl.col("timestamp").dt.minute() % 15 == 0) & (pl.col("timestamp").dt.second() == 0)
    )
    # Keep minute-level for asof alignment to coarser cadences
    predictions_minute = predictions_raw
    print(f"  Predictions: {len(predictions_raw):,} (minute), {len(predictions_15m):,} (15m)")
else:
    predictions_raw = predictions_15m = predictions_minute = None

# %% [markdown]
# ### Aligning predictions to target bar frequency
#
# The predictions are minute-level. When rebalancing at hourly cadence, we take the
# **last available prediction** at or before each price bar timestamp via an asof join.
# This is realistic: the portfolio manager uses the most recent signal when the
# rebalance fires.
#
# A backward asof join produces a null only *before* a symbol's series begins, so
# dropping nulls trims the leading edge and nothing else. Everything after a symbol's
# series *ends* reuses its last score for as long as the panel runs, and the coarser the
# cadence the larger a share of the result that is. On this registry's `fwd_dir_15m`
# predictions, eleven of 113 symbols stop reporting mid-sample - ten of them together at
# the 2020-12-18 fold boundary and `UAL` after a single session - and the oldest match
# the join produced was 252 sessions, the whole panel. At 30-minute cadence 19,667 of
# 303,641 aligned rows carried a prediction more than one session old; at four-hour
# cadence, 3,593 of 55,355.
#
# So the age is bounded, and in sessions rather than minutes: an overnight or weekend
# carry is legitimate and is 5,600 minutes wide, while a symbol that has left the
# universe is months of *sessions* stale. Measured on the same predictions, age 0 and
# age 1 cover every symbol that stays; every row beyond that belongs to one of the
# twelve that leave or have a hole. `MAX_PREDICTION_AGE_SESSIONS` is the bound, and the
# rows it removes are counted and printed rather than dropped quietly - an alignment
# that silently discards a symbol is the same absence this bound exists to end.

# %% [markdown]
# The cadences swept come from `backtest.sweep.cadence_sweep` in `setup.yaml`,
# and the frequency tokens are derived from those names, so the sweep and the
# configuration cannot disagree. An unrecognised cadence stops the notebook
# rather than being skipped, because a silently dropped cadence would leave a
# gap in the heatmap that reads as a region with no viable cost.

# %%
CADENCES = get_cadence_sweep(CASE_STUDY_ID)
_CADENCE_TO_FREQ = {
    "15_minute": "15m",
    "30_minute": "30m",
    "1_hour": "1h",
    "2_hour": "2h",
    "4_hour": "4h",
    "daily_close": "1d",
}
_unknown_cadences = [c for c in CADENCES if c not in _CADENCE_TO_FREQ]
if _unknown_cadences:
    raise ValueError(
        f"cadence_sweep contains unknown cadence(s) {_unknown_cadences!r}; "
        f"valid tokens: {sorted(_CADENCE_TO_FREQ)}"
    )
CADENCE_LABELS = {c: _CADENCE_TO_FREQ[c] for c in CADENCES}
FREQ_MAP = dict(CADENCE_LABELS)

# Per-share cost grid: half-spread + commission in dollars per share.
# Single source of truth is ``backtest.sweep.cost_grid_half_spread_usd`` in
# setup.yaml; labels are derived from the grid so they always match.
COST_PER_SHARE_GRID = get_cost_grid_half_spread_usd(CASE_STUDY_ID)
COST_LABELS = [f"{v * 100:g}¢" for v in COST_PER_SHARE_GRID]

cadence_results = []


# One session, so a bar may use the session's own prediction or the one before it -
# which is what an overnight or weekend gap produces - and nothing older.
MAX_PREDICTION_AGE_SESSIONS = 1


def align_predictions_to_bars(
    preds: pl.DataFrame,
    bar_timestamps: pl.Series,
    *,
    max_age_sessions: int = MAX_PREDICTION_AGE_SESSIONS,
    label: str = "",
) -> pl.DataFrame:
    """Align minute predictions to coarser bars, refusing a match older than the bound.

    Sessions come from the prediction panel's own dates, so the bound counts trading days
    rather than calendar time: a Friday prediction matched to a Monday bar is one session
    old, not three days.

    Returns the aligned frame and, when the bound removed rows, the declaration that has to
    reach `backtest_hash` with them. Without that second value the fix would be invisible to
    the registry: `prediction_hash` and the strategy spec are unchanged by a filter on the
    aligned frame, so a bounded run and an unbounded one hash alike and the later one is
    served the earlier one's result - the same shape as ml4t/agent-workspace#911.
    """
    bar_df = pl.DataFrame({"timestamp": bar_timestamps}).unique().sort("timestamp")
    sessions = preds.select(
        pl.col("timestamp").dt.date().unique().sort().alias("session")
    ).with_row_index("session_index")
    symbols = preds["symbol"].unique().sort().to_list()

    aligned = []
    for sym in symbols:
        sym_preds = preds.filter(pl.col("symbol") == sym).sort("timestamp")
        sym_bars = bar_df.with_columns(pl.lit(sym).alias("symbol"))
        joined = sym_bars.join_asof(
            sym_preds.drop("symbol").with_columns(pl.col("timestamp").alias("prediction_ts")),
            on="timestamp",
            strategy="backward",
        )
        aligned.append(joined.drop_nulls("y_score"))

    if not aligned:
        return pl.DataFrame(), None

    matched = pl.concat(aligned)
    dated = (
        matched.with_columns(pl.col("timestamp").dt.date().alias("session"))
        .join(sessions, on="session", how="left")
        .drop("session")
        .with_columns(pl.col("prediction_ts").dt.date().alias("session"))
        .join(sessions, on="session", how="left", suffix="_prediction")
        .drop("session")
        .with_columns(
            (pl.col("session_index") - pl.col("session_index_prediction")).alias("age_sessions")
        )
    )
    fresh = dated.filter(pl.col("age_sessions") <= max_age_sessions)
    stale = dated.filter(pl.col("age_sessions") > max_age_sessions)
    # Printed on every cadence, including the ones that drop nothing. A line only on the bad
    # case is indistinguishable from the function not having run, which is the shape of the
    # defect this bound closes.
    summary = (
        f"  {label or 'alignment'}: match age {dated['age_sessions'].median():.0f} session(s) "
        f"median, {dated['age_sessions'].max()} max, bound {max_age_sessions}"
    )
    declaration = None
    if stale.height:
        names = sorted(stale["symbol"].unique().to_list())
        summary += (
            f" - dropped {stale.height:,} of {dated.height:,} on "
            f"{len(names)} of {dated['symbol'].n_unique()} symbols "
            f"({', '.join(names[:6])}{' ...' if len(names) > 6 else ''})"
        )
        # Declared only when it removed something, so a cadence the bound does not touch
        # produces the spec it produced before this existed and keeps its registered identity.
        declaration = prediction_age_declaration(
            max_age_sessions=max_age_sessions,
            dropped=stale.height,
            kept=fresh.height,
            symbols_dropped=len(names),
        )
    else:
        summary += " - nothing dropped"
    print(summary)
    aligned = fresh.drop(
        "session_index", "session_index_prediction", "age_sessions", "prediction_ts"
    )
    return aligned, declaration


# %% [markdown]
# ### Run one cadence × cost backtest
#
# Helper that builds the per-share cost spec and runs a single cadence backtest.
# Results are appended to `cadence_results` for the heatmap below.


# %%
def run_cadence_cost_backtest(
    cadence, cadence_label, cost_ps, cadence_prices, aligned_preds, state, prediction_age=None
):
    """Run one cadence × cost backtest and record results."""
    state["n_done"] += 1
    n_done = state["n_done"]

    spec = build_backtest_spec(
        CASE_STUDY_ID,
        bt_config,
        prices=cadence_prices,
        prediction_hash=best_pred_hash,
        initial_cash=bt_config.initial_cash,
        chapter="ch18",
        label=LABEL,
        # `MAX_SYMBOLS` reduced `cadence_prices` and, until the run said so in its own
        # specification, that reduction did not reach `backtest_hash`: a reduced run and the
        # full run over the same predictions hashed alike (ml4t/agent-workspace#911). Built
        # from the panel this spec is being built against, which is the one `run_backtest`
        # is handed below. A full run declares nothing and hashes as it did before.
        traded_universe=(traded_universe_declaration(cadence_prices) if MAX_SYMBOLS else None),
        # Travels with the spec for the same reason the universe does: it changes what this
        # run is computed from and `prediction_hash` cannot see it.
        prediction_age=prediction_age,
        # The universe travels with the spec, not just with the query above. A row registered
        # without it reads as full-universe to every later reader - including section 4's
        # full-versus-screened query and `derived_tables_off_canonical_universe` - so the
        # cadence rows would be filed against the comparison they are not part of.
        signal={
            "method": "equal_weight_top_k",
            "top_k": 20,
            "long_short": bt_config.long_short,
            **({} if CANONICAL_UNIVERSE is None else {"universe_filter": CANONICAL_UNIVERSE}),
        },
    )
    spec["strategy"]["rebalance"]["cadence"] = cadence
    spec["backtest_config"]["metadata"]["cadence"] = cadence

    if cost_ps > 0:
        spec["backtest_config"]["commission"]["model"] = "per_share"
        spec["backtest_config"]["commission"]["per_share"] = cost_ps / 2
        spec["backtest_config"]["commission"]["rate"] = 0.0
        spec["backtest_config"]["slippage"]["model"] = "fixed"
        spec["backtest_config"]["slippage"]["fixed"] = cost_ps / 2
        spec["backtest_config"]["slippage"]["rate"] = 0.0
    else:
        set_backtest_costs_bps(spec, commission_bps=0.0, slippage_bps=0.0)

    spec["cadence_sweep"] = True

    try:
        result = run_backtest(
            CASE_STUDY_ID,
            best_pred_hash,
            spec,
            prices=cadence_prices,
            predictions=aligned_preds,
            label=LABEL,
            register=True,
            initial_cash=bt_config.initial_cash,
            calendar=bt_config.calendar,
        )
        sharpe = result.metrics.get("sharpe", 0)
        n_trades = result.metrics.get("num_trades", 0)

        cadence_results.append(
            {
                "cadence": cadence_label,
                "cost_per_share": cost_ps,
                "cost_label": COST_LABELS[COST_PER_SHARE_GRID.index(cost_ps)],
                "sharpe": sharpe,
                "num_trades": n_trades,
                "cagr": result.metrics.get("cagr", 0),
                "max_drawdown": result.metrics.get("max_drawdown", 0),
            }
        )
        print(
            f"  [{n_done}/{state['n_total']}] {cadence_label} @ {cost_ps * 100:.1f}¢/sh: "
            f"Sharpe={sharpe:.3f}, trades={n_trades:,}"
        )
    except Exception as e:
        print(
            f"  [{n_done}/{state['n_total']}] {cadence_label} @ {cost_ps * 100:.1f}¢/sh: FAILED — {e}"
        )


# %%
sweep_state = {
    "n_total": len(CADENCES) * len(COST_PER_SHARE_GRID) if best_pred_hash else 0,
    "n_done": 0,
}
t0 = time.time()

for cadence in CADENCES if best_pred_hash else []:
    freq = FREQ_MAP[cadence]
    cadence_label = CADENCE_LABELS[cadence]

    cadence_prices = load_backtest_prices_for(
        CASE_STUDY_ID,
        LABEL,
        split="validation",
        frequency=freq,
        max_symbols=MAX_SYMBOLS,
    )
    bar_ts = cadence_prices["timestamp"].unique().sort()

    if freq == "15m":
        # The 15-minute cadence is the prediction grid itself, so no as-of match is made and
        # there is no age to bound.
        aligned_preds, prediction_age = predictions_15m, None
    else:
        aligned_preds, prediction_age = align_predictions_to_bars(
            predictions_minute, bar_ts, label=cadence_label
        )
    if aligned_preds.is_empty():
        print(f"  {cadence_label}: no aligned predictions — skipping")
        continue

    print(
        f"\n--- {cadence_label} cadence: {len(bar_ts)} bars, {len(aligned_preds)} aligned predictions ---"
    )
    for cost_ps in COST_PER_SHARE_GRID:
        run_cadence_cost_backtest(
            cadence,
            cadence_label,
            cost_ps,
            cadence_prices,
            aligned_preds,
            sweep_state,
            prediction_age=prediction_age,
        )

# %%
elapsed_cadence = time.time() - t0
print(f"Cadence sweep: {sweep_state['n_done']} backtests in {elapsed_cadence:.0f}s")

# %% [markdown]
# ### Cadence × Cost Heatmap
#
# This is the central finding: the same signal that is worthless at 15-minute
# cadence becomes viable at hourly cadence with institutional-quality execution
# ($\leq$ 2¢/share effective spread). The table shows Sharpe ratio at each
# cadence × cost combination.

# %%
import matplotlib.pyplot as plt
import numpy as np

cadence_df = pl.DataFrame(cadence_results) if cadence_results else pl.DataFrame()

if not cadence_df.is_empty():
    pivot = cadence_df.pivot(on="cost_label", index="cadence", values="sharpe")
    cadence_order = ["15m", "30m", "1h", "4h"]
    cadences_present = [c for c in cadence_order if c in pivot["cadence"].to_list()]
    costs_present = [c for c in COST_LABELS if c in pivot.columns]

    matrix = np.zeros((len(cadences_present), len(costs_present)))
    for i, cad in enumerate(cadences_present):
        row = pivot.filter(pl.col("cadence") == cad)
        for j, cost_col in enumerate(costs_present):
            if cost_col in row.columns:
                val = row[cost_col][0]
                matrix[i, j] = val if val is not None else np.nan

    fig, ax = plt.subplots(figsize=(10, 5))
    im = ax.imshow(matrix, cmap="RdYlGn", aspect="auto", vmin=-2, vmax=2)
    ax.set_xticks(range(len(costs_present)))
    ax.set_xticklabels(costs_present)
    ax.set_yticks(range(len(cadences_present)))
    ax.set_yticklabels(cadences_present)
    ax.set_xlabel("Effective Spread (per share)")
    ax.set_ylabel("Rebalancing Cadence")
    ax.set_title("Sharpe Ratio: Cadence × Per-Share Cost")

    for i in range(len(cadences_present)):
        for j in range(len(costs_present)):
            val = matrix[i, j]
            color = "white" if abs(val) > 1.0 else "black"
            ax.text(j, i, f"{val:.2f}", ha="center", va="center", color=color, fontsize=11)

    fig.colorbar(im, ax=ax, label="Sharpe Ratio")
    show_with_alt(
        fig,
        "Heatmap of Sharpe over rebalancing cadence against per-share effective spread. "
        "Rows are cadences and columns are spreads, each cell holding one Sharpe printed to "
        "two decimals over a red-yellow-green colour scale clipped at -2 and +2, so red is "
        "the worst and green the best. A colour bar to the right carries the same scale. The "
        "cell values are the figure's data: the colour repeats them rather than adding "
        "anything a reader would otherwise miss.",
    )
else:
    print("No cadence sweep results")

# %%
if not cadence_df.is_empty():
    print("=== Cadence × Cost Summary ===")
    print(
        cadence_df.sort("cadence", "cost_per_share").select(
            "cadence", "cost_label", "sharpe", "num_trades"
        )
    )

# %% [markdown]
# ### Trade Count by Cadence
#
# Reducing the rebalancing cadence cuts trade counts dramatically, which is
# the mechanism behind the Sharpe improvement: fewer trades means less
# cumulative cost drag. The trade-off is signal decay — the 15-minute
# prediction becomes stale at longer horizons. The sweet spot for this
# dataset is hourly cadence where the signal retains enough edge to cover
# 1–2¢/share execution costs.

# %%
if not cadence_df.is_empty():
    zero_cost = cadence_df.filter(pl.col("cost_per_share") == 0.0)
    if not zero_cost.is_empty():
        fig, axes = plt.subplots(1, 2, figsize=(12, 5))

        cadences = zero_cost["cadence"].to_list()
        trades = zero_cost["num_trades"].to_list()
        sharpes = zero_cost["sharpe"].to_list()

        axes[0].barh(cadences, trades)
        axes[0].set_xlabel("Number of Trades")
        axes[0].set_title("Trade Count by Cadence (Zero Cost)")

        axes[1].barh(cadences, sharpes)
        axes[1].axvline(0, color="gray", linestyle="--")
        axes[1].set_xlabel("Sharpe Ratio")
        axes[1].set_title("Gross Sharpe by Cadence")

        show_with_alt(
            fig,
            "Two horizontal bar charts side by side, sharing one cadence per row. The left "
            "panel gives the number of trades each cadence closes and the right panel its "
            "gross Sharpe, both at zero assumed cost, with a dashed vertical line at zero "
            "Sharpe on the right. Reading a row across both panels is what the pair is for: "
            "it puts the trading a cadence provokes beside the return it earns before any "
            "cost is charged.",
        )

# %% [markdown]
# ## Key Takeaways
#
# 1. **Costs are charged per trade, so trading frequency sets what they take.**
#    A cost that is negligible against a multi-day return is decisive against a
#    fifteen-minute one, and the same strategy can be profitable or not on that
#    difference alone. Any strategy result quoted without its cost assumption is
#    incomplete.
#
# 2. **Restricting the universe and slowing the cadence are different levers.**
#    One changes what each trade costs and which names the ordering can act on;
#    the other changes how many trades there are. They are measured separately
#    here because a result that changed both at once cannot attribute its
#    improvement to either.
#
# 3. **A screen on tradability is not a screen on the signal.** Keeping the
#    cheapest-to-trade names uses no information about returns, and it is fixed
#    per split so it cannot see the window it is evaluated on. That is what
#    keeps it a cost decision rather than a selection.
#
# 4. **The breakeven cost is the number to carry forward.** It states the
#    execution quality a strategy requires rather than the profit it produced
#    under one assumption, and it can be checked against what a given venue and
#    order size actually achieve.
#
# **Known limitations**: The cadence sweep splits its cost evenly between the
# engine's commission and slippage settings so a single dollars-per-share axis
# can be swept. Production costs use a fixed per-share commission plus a
# measured per-asset half-spread, which are different quantities on different
# scales. The sweep is therefore a co

Полный текст с указанием источника опубликован на условиях его лицензии. Лицензия: MIT

Это краткое изложение подготовлено исследовательским агентом Stratmill по оригиналу и не является его копией.