跳至正文
返回文库全部文档

评估交易成本下的策略存续能力

笔记本 《交易机器学习》

总结

本笔记本比较多个案例策略在假设交易成本下的影响。它读取每项已部署策略配置的成本敏感度回测,将毛夏普率与案例自身假设成本下的净夏普率进行比较,并估算夏普率降至零时的成本。如果假设成本落在已测试成本点之间,则对净夏普率进行线性插值。分析还按调仓周期分组,考察换手频率与成本拖累的关系。

笔记本强调如何解读缺失曲线和受限的盈亏平衡估计:若策略在已测试成本最高点仍为正,其盈亏平衡成本可能高于该范围,但具体值未知。文中还解释,若主要成本来自较宽点差,通用的比例基点扫描可能无法准确代表成本;这时基于报价的可执行性测试更合适。证据仅限于已部署策略中可用的成本扫描,使用假设的周期换手乘数,并依赖在验证数据上筛选出的配置所对应的验证结果。实际成本也会随规模、工具和市场状况变化。

核心观点

  • 在各策略自身假设的交易成本下比较毛夏普率和净夏普率。
  • 在成本网格点之间插值估算净表现,并报告所用区间。
  • 将超出测试范围的盈亏平衡值视为下限,而不是已测得的交叉点。
  • 基于周期的换手乘数是人为假设,并非从回测中测得的换手率。
  • 点差结构占主导时,比例基点扫描可能会误估成本。

标签

全文
# Friction Survival: Where the Edge Dies


# Friction Survival: Where the Edge Dies

**Docker image**: `ml4t`

This notebook cross-cuts the case studies by **failure mode**: gross-to-net
Sharpe degradation, breakeven cost thresholds, cadence-frequency vulnerability,
and cost-model realism caveats. Ch18 establishes the cost taxonomy and the
per-asset-class machinery; this notebook reads the resulting Ch18 cost-sweep
backtests directly out of each case study's registry and asks which
strategies survive the friction of real trading.

**Learning Objectives**:
- Quantify each strategy's sensitivity to per-leg friction on a common scale
- Identify breakeven cost thresholds per rebalance cadence
- Recognize when a generic basis-point sweep is the wrong cost model
  (single-name options, intraday equities)

**Book Reference**: Chapter 20, Section 20.6 (Trading Realism)

**Prerequisites**: Run [`01_aggregate_synthesis`](01_aggregate_synthesis.ipynb) first.
Each case study's registry must contain Ch18 `cost_sensitivity`-stage backtests.

```python
"""Ch20 Friction Survival — cross-case-study cost-sweep analysis from registry."""

import json

import matplotlib.pyplot as plt
import polars as pl
from IPython.display import Markdown, display
from matplotlib.patches import Patch

from case_studies.utils.analytics import (
    CASE_STUDY_IDS,
    SHORT_NAMES,
    load_carrier_cost_curves,
)
from utils.paths import get_chapter_dir
from utils.style import show_with_alt

pl.Config.set_tbl_rows(20)
```

```python
# 0 = all
MAX_CASE_STUDIES = 0
```

```python
CS_LIST = CASE_STUDY_IDS[:MAX_CASE_STUDIES] if MAX_CASE_STUDIES else CASE_STUDY_IDS
```

## Load Cost Sweep Results from Registry

Ch18 backtests vary commission + slippage across a grid of cost levels
while holding the signal and allocation constant. We read the sweep for
each case study's **release configuration** -- the one declared across
the signal, allocation, and risk-overlay stages --
so the breakeven measured here is the cost survival of the strategy the
chapter actually deploys, not of whichever allocator happened to be best
at zero cost.

A case study can hold cost-sensitivity backtests and still draw no curve here, because
the sweep has to sit on the carrier's own training lineage *and* run the carrier's own
strategy. The loader reports which check dropped each one, printed below the load, so an
absence from the charts can be read rather than guessed at.

```python
loaded = load_carrier_cost_curves(CS_LIST)
costs_df = loaded.curves
_exclusions = loaded.exclusion_lines()

if costs_df.is_empty():
    # The reasons go into the refusal rather than after it. Every case study being excluded is
    # the state that most needs them - a clean clone with no registries reaches it - and the
    # loader has already established each one.
    msg = "No Ch18 cost-sensitivity backtests found for any deployed carrier"
    raise RuntimeError("\n".join([msg, *_exclusions]) if _exclusions else msg)

n_cs = costs_df["case_study"].n_unique()
print(f"Loaded {len(costs_df)} carrier cost-sweep entries across {n_cs} case studies")
for _line in _exclusions:
    print(_line)
costs_df.head(5)
```

Assumed per-leg cost and rebalance cadence both come from each case study's
`setup.yaml`, read back through the Ch20 artifacts that `01_aggregate_synthesis`
writes, rather than being typed into this notebook.

`setup.yaml` records a specific cadence such as `monthly_month_end` or
`daily_ny_close`. The charts group by period, taken from the leading word, and
anything unrecognised is grouped as unspecified rather than given a default.
The turnover multipliers attached to each period say how often a book of that
cadence is assumed to turn over relative to a daily one. They are an
assumption, not turnover measured from the backtests.

## Gross-to-Net Sharpe Degradation

For each case study we compare the zero-cost (gross) Sharpe with the Sharpe at
the cost that case study actually assumes, which comes from its `setup.yaml`
by way of `overview.parquet`. The assumed cost rarely falls on a grid point, so
the net Sharpe is interpolated linearly between the two grid points that
bracket it, and the bracketing points are reported alongside.


```python
gross_df = costs_df.filter(pl.col("cost_bps") == 0)
net_df = costs_df.filter(pl.col("cost_bps") > 0)

# One allocator per case study (the selected configuration's); this selects it.
best_alloc = (
    gross_df.sort("sharpe", descending=True)
    .unique(subset=["case_study"], keep="first")
    .select("case_study", "allocator")
)
```

```python
_overview = pl.read_parquet(get_chapter_dir(20) / "output" / "overview.parquet")
ASSUMED_COST_BPS = dict(_overview.select("cs_id", "cost_bps").iter_rows())
_synthesis = json.loads((get_chapter_dir(20) / "output" / "all_synthesis.json").read_text())
CADENCE_BY_CS = {
    cs: (data["meta"].get("cadence") or "unspecified") for cs, data in _synthesis.items()
}

CADENCE_PERIODS = ("15min", "hourly", "8_hour", "daily", "weekly", "monthly")
TURNOVER_MULTIPLIER = {
    "15min": 26.0,
    "hourly": 6.5,
    "8_hour": 3.0,
    "daily": 1.0,
    "weekly": 1.0 / 5,
    "monthly": 1.0 / 21,
}


def _cadence_period(cadence: str) -> str:
    for period in CADENCE_PERIODS:
        if cadence.startswith(period):
            return period
    return "unspecified"


def _sharpe_at(curve: pl.DataFrame, cost_bps: float) -> tuple[float, float, float]:
    """Sharpe at `cost_bps`, linearly interpolated on the sweep grid.

    Returns the interpolated Sharpe and the two grid costs it sits between. A
    cost beyond either end of the grid is clamped to that end, which is reported
    by the bracket coming back equal.
    """
    grid = curve["cost_bps"].to_list()
    vals = curve["sharpe"].to_list()
    if cost_bps <= grid[0]:
        return vals[0], grid[0], grid[0]
    if cost_bps >= grid[-1]:
        return vals[-1], grid[-1], grid[-1]
    for lo, hi, v_lo, v_hi in zip(grid, grid[1:], vals, vals[1:], strict=False):
        if lo <= cost_bps <= hi:
            w = 0.0 if hi == lo else (cost_bps - lo) / (hi - lo)
            return v_lo + w * (v_hi - v_lo), lo, hi
    return vals[-1], grid[-1], grid[-1]


def _breakeven(curve: pl.DataFrame) -> tuple[float, bool]:
    """Cost at which Sharpe crosses zero, and whether that crossing was observed.

    Interpolates between the last positive grid point and the first negative one.
    A curve still positive at the top of the grid is censored: the breakeven is
    somewhere above the ceiling and the ceiling is not it.
    """
    grid = curve["cost_bps"].to_list()
    vals = curve["sharpe"].to_list()
    for lo, hi, v_lo, v_hi in zip(grid, grid[1:], vals, vals[1:], strict=False):
        if v_lo > 0 >= v_hi:
            w = v_lo / (v_lo - v_hi)
            return lo + w * (hi - lo), True
    return (grid[-1], False) if vals[-1] > 0 else (0.0, True)


summary_rows = []
for row in best_alloc.iter_rows(named=True):
    cs = row["case_study"]
    alloc = row["allocator"]
    curve = costs_df.filter((pl.col("case_study") == cs) & (pl.col("allocator") == alloc)).sort(
        "cost_bps"
    )
    if curve.height < 2 or curve["cost_bps"].min() > 0:
        continue

    gross_sharpe = curve["sharpe"][0]
    assumed_cost = ASSUMED_COST_BPS.get(cs)
    if assumed_cost is None:
        msg = f"No cost_bps for {cs} in overview.parquet; re-run 01_aggregate_synthesis"
        raise RuntimeError(msg)
    net_sharpe, bracket_lo, bracket_hi = _sharpe_at(curve, assumed_cost)
    breakeven_bps, breakeven_observed = _breakeven(curve)

    summary_rows.append(
        {
            "case_study": cs,
            "display_name": SHORT_NAMES.get(cs, cs),
            "cadence": CADENCE_BY_CS.get(cs, "unspecified"),
            "cadence_period": _cadence_period(CADENCE_BY_CS.get(cs, "unspecified")),
            "allocator": alloc,
            "gross_sharpe": round(gross_sharpe, 3),
            "net_sharpe": round(net_sharpe, 3),
            "sharpe_drag": round(gross_sharpe - net_sharpe, 3),
            "drag_pct": round(100 * (gross_sharpe - net_sharpe) / gross_sharpe, 1)
            if gross_sharpe != 0
            else 0.0,
            "assumed_cost_bps": assumed_cost,
            "grid_bracket": f"{bracket_lo:g}-{bracket_hi:g}",
            "breakeven_bps": round(breakeven_bps, 1),
            "breakeven_observed": breakeven_observed,
            "survives": net_sharpe > 0,
        }
    )
```

```python
summary = pl.DataFrame(summary_rows).sort("drag_pct", descending=True)
print("=== Gross-to-Net Sharpe Degradation ===")
summary.select(
    "display_name",
    "cadence",
    "gross_sharpe",
    "assumed_cost_bps",
    "grid_bracket",
    "net_sharpe",
    "sharpe_drag",
    "drag_pct",
    "breakeven_bps",
    "breakeven_observed",
    "survives",
)
```

```python
_dead = summary.filter(~pl.col("survives"))
_censored = summary.filter(~pl.col("breakeven_observed"))
display(
    Markdown(
        f"{summary.height} case studies have a carrier cost sweep. "
        + (
            f"**{', '.join(_dead['display_name'].to_list())}** "
            f"{'has' if _dead.height == 1 else 'have'} a negative Sharpe at the "
            "cost the case study assumes, so the strategy does not survive its "
            "own cost model. "
            if _dead.height
            else "All of them keep a positive Sharpe at the cost they assume. "
        )
        + f"Cost consumes between {summary['drag_pct'].min():.1f} and "
        f"{summary['drag_pct'].max():.1f} percent of gross Sharpe.\n\n"
        + (
            f"For {', '.join(_censored['display_name'].to_list())} the Sharpe is "
            f"still positive at the top of the swept grid, so the breakeven "
            "column is a lower bound rather than a measurement: it is at least "
            "that, and the grid does not say how much more."
            if _censored.height
            else "Every breakeven was observed inside the swept grid."
        )
    )
)
```

## Cost Drag Visualization

The horizontal bar chart shows Sharpe drag (gross minus net) for each
case study, ordered by severity. Higher-frequency strategies typically
suffer more because they accumulate turnover costs faster.

```python
fig, ax = plt.subplots(figsize=(10, 6))

colors = [
    "#d62728" if drag > 50 else "#ff7f0e" if drag > 20 else "#2ca02c"
    for drag in summary["drag_pct"]
]

bars = ax.barh(
    range(len(summary)),
    summary["drag_pct"].to_list(),
    color=colors,
    edgecolor="none",
    height=0.6,
)
ax.set_yticks(range(len(summary)))
ax.set_yticklabels(summary["display_name"].to_list())
ax.set_xlabel("Sharpe Drag (%)")
ax.set_title("Cost Impact: Gross-to-Net Sharpe Degradation")
ax.invert_yaxis()

for bar, row in zip(bars, summary.iter_rows(named=True), strict=False):
    ax.annotate(
        f"BE: {'' if row['breakeven_observed'] else '>'}{row['breakeven_bps']:g} bps",
        xy=(bar.get_width() + 1, bar.get_y() + bar.get_height() / 2),
        va="center",
        fontsize=9,
        color="gray",
    )
# Headroom so the breakeven annotation on the widest bar (FX) is not clipped.
ax.set_xlim(right=max(summary["drag_pct"]) * 1.28)

show_with_alt(
    fig,
    "Horizontal bars giving the percentage of gross Sharpe consumed by each case "
    "study's assumed cost, ordered by severity, each annotated with the cost at "
    "which that strategy breaks even.",
)
```

## Breakeven Cost Thresholds by Frequency

Breakeven cost is the maximum per-leg cost (in bps) at which the
deployed configuration still produces a positive Sharpe ratio. It is the cost
budget that the signal supports before becoming unprofitable.

```python
freq_order = list(CADENCE_PERIODS)
freq_colors = {
    "15min": "#d62728",
    "hourly": "#e8833a",
    "8_hour": "#ff7f0e",
    "daily": "#1f77b4",
    "weekly": "#5aa469",
    "monthly": "#2ca02c",
}

fig, ax = plt.subplots(figsize=(10, 5))

for i, row in enumerate(summary.sort("breakeven_bps").iter_rows(named=True)):
    color = freq_colors.get(row["cadence_period"], "gray")
    ax.barh(i, row["breakeven_bps"], color=color, height=0.6, edgecolor="none")

ax.set_yticks(range(len(summary)))
sorted_names = summary.sort("breakeven_bps")["display_name"].to_list()
ax.set_yticklabels(sorted_names)
ax.set_xlabel("Breakeven Cost (bps per leg)")
ax.set_title("Breakeven Cost Thresholds — Higher Is More Robust")

legend_handles = [Patch(facecolor=freq_colors[f], label=f) for f in freq_order if f in freq_colors]
ax.legend(handles=legend_handles, loc="lower right", title="Cadence")

show_with_alt(
    fig,
    "Horizontal bars of the breakeven per-leg cost for each case study, ordered "
    "from lowest to highest and coloured by rebalance cadence.",
)
```

## Cost Drag Curves

For each case study, plot Sharpe ratio as a function of per-leg
cost. This reveals the "cost cliff" — the point where a profitable
strategy becomes unprofitable.

```python
best_alloc_map = dict(
    zip(best_alloc["case_study"].to_list(), best_alloc["allocator"].to_list(), strict=False)
)

fig, ax = plt.subplots(figsize=(12, 7))

for cs_id in CS_LIST:
    alloc = best_alloc_map.get(cs_id)
    if alloc is None:
        continue

    cs_data = costs_df.filter(
        (pl.col("case_study") == cs_id) & (pl.col("allocator") == alloc)
    ).sort("cost_bps")

    if cs_data.is_empty():
        continue

    ax.plot(
        cs_data["cost_bps"].to_list(),
        cs_data["sharpe"].to_list(),
        marker="o",
        markersize=4,
        label=SHORT_NAMES.get(cs_id, cs_id),
    )

ax.axhline(y=0, color="black", linestyle="--", alpha=0.3, linewidth=0.8)
ax.set_xlabel("Per-Leg Cost (bps)")
ax.set_ylabel("Sharpe Ratio")
ax.set_title("Cost sensitivity: Sharpe against per-leg cost")
ax.legend(loc="upper right", fontsize=9, ncol=2)

# Mark each case study's assumed cost so the curve can be read at the point that
# matters rather than across the whole grid.
for cs_id in best_alloc_map:
    _c = ASSUMED_COST_BPS.get(cs_id)
    if _c is not None:
        ax.axvline(_c, color="gray", alpha=0.25, linewidth=0.8, linestyle=":")

show_with_alt(
    fig,
    "Line chart of Sharpe against per-leg cost in basis points, one line per "
    "case study over the swept grid, with a reference line at zero Sharpe and "
    "faint vertical lines marking each case study's assumed cost.",
)
```

## Cost Survival Classification

Each case study is classified by cost resilience: the ratio of its breakeven to
the per-leg cost it is assumed to pay. A higher ratio means more headroom once
realistic frictions are imposed. A ratio below one means the breakeven sits
under the assumed cost, so the strategy is already losing money at its own
assumption, and it is classified apart from a thin but positive margin.

The assumed cost is the one already in `summary`, read from each case study's
own setup rather than declared again here, so this table and the degradation
table above cannot disagree about what a case study is assumed to pay.

```python
survival = summary.with_columns(
    cost_margin_bps=(pl.col("breakeven_bps") - pl.col("assumed_cost_bps")),
    cost_margin_ratio=(pl.col("breakeven_bps") / pl.col("assumed_cost_bps").clip(lower_bound=1)),
).with_columns(
    resilience=pl.when(pl.col("cost_margin_ratio") < 1)
    .then(pl.lit("does not survive"))
    .when(pl.col("cost_margin_ratio") >= 10)
    .then(pl.lit("very robust"))
    .when(pl.col("cost_margin_ratio") >= 3)
    .then(pl.lit("robust"))
    .when(pl.col("cost_margin_ratio") >= 1.5)
    .then(pl.lit("marginal"))
    .otherwise(pl.lit("fragile")),
)

print("=== Cost Survival Classification ===")
survival.select(
    "display_name",
    "cadence",
    "assumed_cost_bps",
    "net_sharpe",
    "breakeven_bps",
    "breakeven_observed",
    "cost_margin_ratio",
    "resilience",
)
```

```python
_res = survival.group_by("resilience").agg(cs=pl.col("display_name")).sort("resilience")
display(
    Markdown(
        "; ".join(
            f"**{r['resilience']}**: {', '.join(sorted(r['cs']))}"
            for r in _res.iter_rows(named=True)
        )
        + ". A ratio is only as good as the breakeven behind it, and where the "
        "sweep never crossed zero the breakeven is the grid ceiling rather than "
        "a crossing, so the ratio for those is a lower bound too."
    )
)
```

## S&P 500 Options: Spread Realism Caveat

The S&P 500 Options case study was validated using executable-label
backtesting, pricing straddle entries and exits at actual bid/ask quotes rather
than at an assumed bps cost. That case study has no selected configuration cost sweep, so it
does not appear in any table above.

It is described here for the structure of its cost problem rather than for its numbers, which
its own evaluation and §18.8 carry. A single-name option's dominant execution cost is the
bid-ask spread on the premium rather than a commission proportional to notional, so the cost
scales with how wide the quote is and not with how much is traded. That is why its evaluation
decomposes one prediction across three labels - priced at the mid and unhedged, delta-hedged
at the mid, and priced at the quotes a desk would actually get - which separates the signal's
contribution from the execution's, and why ranking on signal and spread jointly is a different
strategy from ranking on signal alone rather than a refinement of it.

A generic bps cost sweep misrepresents this case study for the same reason: it models a cost
that is proportional to notional. The teaching point is that strategy design has to optimize
for signal quality and execution cost together, because for this instrument the spread is what
the signal has to pay for.

## Cadence–Frequency–Cost Regime

The same IC translates to very different tradability depending on
rebalance cadence. A 15-minute strategy accumulates ~25× more turnover
per day than a daily strategy, and ~500× more than a monthly one.
This creates distinct cost regimes:

```python
if not summary.is_empty():
    regime = summary.with_columns(
        turnover_mult=pl.col("cadence_period").replace_strict(
            TURNOVER_MULTIPLIER,
            default=1.0,
            return_dtype=pl.Float64,
        ),
    )
```

```python
if not summary.is_empty():
    fig, ax = plt.subplots(figsize=(10, 6.5))

    # Turnover-mult on x (varies 0.05→26×); breakeven on y. Both log so the
    # high-frequency cluster (NQ100/Crypto) and the monthly cluster
    # separate cleanly instead of stacking on a constant-x degenerate column.
    assumed_floor = max(float(summary["assumed_cost_bps"].min()), 0.5)

    # Monthly selected configurations share x (turnover ≈ 0.05) and pair up on y: ETFs and
    # US Firms at 50, CME and SP500 Eq+Opt at 30. Fan their labels vertically
    # so the two pairs stay legible despite the superimposed markers.
    label_offsets = {
        "NQ100": (10, 4),
        "Crypto": (10, 4),
        "FX": (10, 4),
        "US Equities": (10, 4),
        "ETFs": (10, 16),
        "US Firms": (10, 2),
        "SP500 Eq+Opt": (10, -2),
        "CME Futures": (10, -16),
        "SP500 Options": (10, 4),
    }

    for row in regime.iter_rows(named=True):
        color = freq_colors.get(row["cadence_period"], "gray")
        size = max(60, min(360, row["turnover_mult"] ** 0.5 * 120))
        ax.scatter(
            row["turnover_mult"],
            max(row["breakeven_bps"], 0.5),
            s=size,
            c=color,
            edgecolors="white",
            linewidth=1.2,
            zorder=5,
        )
        dx, dy = label_offsets.get(row["display_name"], (8, 8))
        ax.annotate(
            row["display_name"],
            (row["turnover_mult"], max(row["breakeven_bps"], 0.5)),
            xytext=(dx, dy),
            textcoords="offset points",
            fontsize=9,
            zorder=6,
        )

    ax.axhline(
        assumed_floor,
        color="0.35",
        linestyle="--",
        linewidth=1.0,
        zorder=3,
        label=f"Survival floor ({assumed_floor:.0f} bps assumed cost)",
    )
    ax.set_xscale("log")
    ax.set_yscale("symlog", linthresh=1)
    ax.set_xlim(0.03, 60)
    ax.set_ylim(-0.5, 600)

    ax.set_xlabel("Turnover multiplier vs daily (log)")
    ax.set_ylabel("Breakeven cost — bps per leg (symlog)")
    ax.set_title("Cost Regimes: Higher-Frequency Strategies Face Steeper Cliffs")

    legend_handles = [
        Patch(facecolor=freq_colors[f], label=f) for f in freq_order if f in freq_colors
    ]
    ax.legend(
        handles=legend_handles + [ax.get_lines()[0]],
        loc="upper right",
        title="Cadence",
        framealpha=0.9,
    )

    show_with_alt(
        fig,
        "Log-log scatter of breakeven cost against assumed relative turnover, one "
        "marker per case study coloured by cadence, with a horizontal line at the "
        "lowest assumed cost in the panel.",
    )
```



```python
_reg = regime.sort("turnover_mult", descending=True)
display(
    Markdown(
        "Marker x-position is assumed per-day turnover relative to a daily "
        "strategy, y is the cost at which net Sharpe crosses zero. The turnover "
        "multipliers are an assumption written into this notebook, not a "
        "measurement from the backtests: they say how often a book of a given "
        "cadence is expected to turn over, and the chart uses them to place the "
        "case studies rather than to test them.\n\n"
        + "; ".join(
            f"**{r['display_name']}** ({r['cadence']}), breakeven "
            f"{'' if r['breakeven_observed'] else 'at least '}"
            f"{r['breakeven_bps']:g} bps against an assumed "
            f"{r['assumed_cost_bps']:g}"
            for r in _reg.iter_rows(named=True)
        )
        + ".\n\nThe cadences present here span a narrow part of the range the "
        "chart is drawn for. The high-frequency corner is empty: NASDAQ-100's "
        "cost sweep ran a different strategy from its carrier, and the other "
        "sub-daily case studies have no cost sweep. Nothing here tests whether "
        "turnover or signal strength sets the breakeven, because the case "
        "studies that would separate them are the ones missing."
    )
)
```

## Key Takeaways

- **A cost sweep is only informative at the cost the strategy assumes.** The
  gross Sharpe and the Sharpe at the top of the grid are both easy to read off
  and neither is the number that decides whether the strategy is tradable. The
  assumed cost comes from the case study's own setup, and the tables above
  report the net Sharpe there.
- **Breakeven and assumed cost have to be compared, not reported side by
  side.** The ratio between them is the headroom, and a ratio below one means
  the strategy is already under water at its own assumption. The computed
  classification above says which case studies are where.
- **A breakeven above the top of the swept grid is not a breakeven.** Where the
  curve is still positive at the ceiling, the honest statement is that the
  crossing is somewhere above it, and the tables mark those rows rather than
  printing the ceiling as though it had been measured.
- **A basis-point grid does not model every cost structure.** Where the
  dominant cost is a wide bid-ask spread rather than a proportional fee, a bps
  sweep understates it, and the answer is an executable backtest against
  quotes. The S&P 500 Options section above is the worked case.

## Known Limitations

- Only case studies whose *carrier* has a cost sweep appear. The loaded count and
  one line per absent case study are printed at the top, so which check dropped a
  case study is read off the run rather than reconstructed by hand. A case study can
  hold cost-sensitivity backtests and still be absent, because the sweep has to sit
  on the deployed carrier's own training lineage and run the carrier's own strategy.
  Three are absent and each fails a different check. ETFs' carrier lineage carries no
  cost sweep at all. S&P 500 Options has eight cost rows on its carrier's lineage,
  all of them an `equal_weight_top_k` + `score_weighted` series rather than the
  carrier. NASDAQ-100 has 24 on its carrier's lineage, all of them `equal_weight_top_k`
  - the instrument its pass-1 ranking uses - while its carrier is a
  `slot_persistent_signal_exit` strategy. All three absences are properties of what
  was swept rather than of this chapter: a carrier with no cost sweep of its own has
  no cost curve to draw.
- The sweep applies one proportional per-leg cost to every trade. Real costs
  vary with size, with the instrument, and with the state of the book, and the
  spread realism section is where that assumption is checked rather than
  assumed.
- Net Sharpe at the assumed cost is interpolated between grid points; the
  bracketing points are in the table so the interpolation can be checked.
- The turnover multipliers used to place case studies on the cadence chart are
  stated assumptions about how often each cadence trades, not turnover measured
  from the backtests.
- Every Sharpe here is a validation-fold number for a configuration chosen on
  validation data, so the cost headroom inherits that selection.

**Next**: [`07_regime_risk`](07_regime_risk.ipynb) examines regime
robustness and risk overlays.
![notebook output](figures/p1_1.png)
![notebook output](figures/p1_2.png)
![notebook output](figures/p1_3.png)
![notebook output](figures/p1_4.png)

在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT

此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。