Saltar al contenido
Todos los documentos de la biblioteca

Medición de la sensibilidad de una estrategia FX a los costes de transacción

Código Machine Learning for Trading

Resumen

Este cuaderno mide cómo responde a los cambios en los costes de transacción proporcionales una estrategia FX seleccionada en validación. Para cada etiqueta de predicción, selecciona una estrategia base entre los resultados de señal, asignación y superposición de riesgo, y luego vuelve a ejecutar esa estrategia en una cuadrícula de costes, conservando los demás campos de identidad. El barrido de costes es un análisis de perturbación: describe la solidez de una estrategia ya definida y no interviene en la selección posterior.

El documento explica por qué los puntos básicos son adecuados para los diferenciales FX, que son proporcionales al tipo de cambio, y por qué el backtest aplica un coste agregado por tramo negociado aunque la configuración enumere por separado los componentes de diferencial y swap. Interpreta la curva mediante la rotación de cartera y su cruce por cero en relación con los niveles cotizados del diferencial. Estos resultados corresponden al periodo de validación, y los supuestos de tasas constantes no pueden representar la ampliación de diferenciales junto con la volatilidad ni los cambios de rotación entre regímenes. El fragmento presenta un método y sus limitaciones, pero no resultados realizados de la estrategia.

Ideas clave

  • Elige la estrategia base entre los candidatos de validación de las etapas de señal, asignación y riesgo antes de probar los costes.
  • Excluye las variantes de costes de la selección posterior para evitar favorecer supuestos con fricciones artificialmente bajas.
  • Representa proporcionalmente las fricciones FX, ya que los diferenciales se expresan en relación con los tipos de cambio.
  • Interpreta la sensibilidad a costes como una relación entre la rotación de cartera, el rendimiento neto y la tasa de coste asumida.
  • Una curva de costes constante del periodo de validación no puede reflejar cambios de rotación y ampliaciones de diferenciales dependientes del régimen.

Etiquetas

Texto completo
# 16_costs.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Transaction-Cost Sensitivity - FX Pairs
#
# This notebook selects one validation strategy per label from the equal-weight, allocation and
# risk-overlay populations, then changes only its percentage transaction costs. The sensitivity
# grid does not participate in later model or strategy selection.
#
# Costs are swept last because a cost curve is only informative about the strategy that would
# actually be traded, and that strategy is not settled until the risk controls have been measured.
# Sweeping before the overlay charges the grid against a configuration the next notebook may
# discard.
#
# This is a perturbation analysis, not a choice. It asks what happens to a settled strategy if the
# cost model is wrong, which is a different question from which strategy to trade. The two must
# stay separate for a mechanical reason: a strategy allowed to compete on its own cost assumption
# would win by having costs assumed away, and the ranking would report the most optimistic
# assumption rather than the best strategy.
#
# **Why basis points here.** FX spreads are quoted in pips, a fraction of the rate itself, so a
# proportional charge is how the friction is actually expressed - there is no share or contract
# to bill per unit. Case studies where nominal prices are stable, or where spreads are measured
# from quote data, use per-share instead; applying a flat per-unit charge to a rate would assume
# spread scales with the level, which it does not. The configured grid runs from 0 to 50 basis
# points per traded leg, against a real quoted band of roughly 1 to 3 for major pairs and 3 to 8
# for crosses, so most of the grid sits deliberately past anything plausible.
#
# `config/setup.yaml` names the cost components as spread and swap points. That list is a
# taxonomy, read by the Chapter 18 teaching notebooks to describe what the friction consists of;
# the backtest charges a single aggregate rate per traded leg. So this curve perturbs the
# aggregate, and the overnight financing cost of carrying a position is described rather than
# separately priced.
#
# **Learning objectives**
#
# - Select from an immutable validation candidate set by backtest Sharpe.
# - Preserve model, checkpoint, signal, allocation, and execution identities across a cost curve.
# - Keep cost sensitivity outside the official selection cohort.
# - Read a cost curve as a statement about turnover.
#
# **Book reference**: Chapter 18
#
# **Prerequisite**: `15_risk_management`.

# %%
"""Run one cost-sensitivity curve per FX prediction label."""

from copy import deepcopy
from typing import Any

import polars as pl
import yaml

from case_studies.research import (
    BacktestResult,
    CandidateSet,
    OfficialPopulation,
    Result,
    candidate_set_supersedes,
    open_study,
    plan_backtests,
    population_supersedes,
    research_name,
    reuse_disclosure,
    run_backtests,
    superseded_members,
)
from case_studies.utils.backtest_presets import EngineBacktestConfig
from case_studies.utils.strategy_analysis import selectable_validation_candidates
from case_studies.utils.sweep_config import (
    get_cost_grid_bps,
)
from utils.paths import get_case_study_dir
from utils.reproducibility import set_global_seeds

# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
SPLIT = "validation"
TOP_K = 0
TOP_N_PREDICTIONS = None
MAX_COST_POINTS = 0
SEED = 42
RUN_SWEEP = True
FORCE_REBACKTEST = False
POPULATION_NAME = ""
SUPERSEDES_COST_BACKTESTS: str = "live"
# The same rule the populations follow: a candidate set is immutable under its name, so a rebuilt
# upstream generation must name the set it replaces. Keyed by the full set name, which is what the
# refusal prints. `15_risk_management` states the reasoning once.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "fx_pairs:fwd_ret_1d:pre-cost-strategies": "live",
    "fx_pairs:fwd_ret_5d:pre-cost-strategies": "live",
    "fx_pairs:fwd_ret_21d:pre-cost-strategies": "live",
}

# %% [markdown]
# ## Select one strategy for each label
#
# Production selection considers the signal, allocation and risk-overlay populations together -
# the same three stages the canonical validation rank-1 is selected over in
# `case_studies/utils/strategy_analysis.py:resolve_canonical_rank1_lineage`. All three are
# candidates because each stage is an alternative to the one before it rather than an improvement
# on it by construction. Naming only the overlays would charge the cost grid against a risk
# control even where every control measured worse than leaving the position rule alone; naming
# only signal and allocation would sweep a strategy the risk notebook has already improved on.
# Which stage the parent came from is printed below rather than assumed, and the gap between the
# best overlay and the best un-overlaid configuration is printed with it: a negative gap is the
# measurement that the controls did not help on this label.
#
# Cost variants are descendants of that choice and cannot improve their own chance of selection.
# Preview mode uses a deterministic allocation request from the reduced catalog and remains
# outside candidate sets.
#
# Getting the parent wrong has a quiet failure mode. The cost curve would be computed correctly,
# the population would freeze and validate, and every number would be right - about a strategy
# the chapter does not report. Nothing raises, because a cost sweep over the wrong parent is a
# perfectly valid sweep. That is why the stage the parent came from is printed rather than
# assumed, and why the selection here is made over the same three stages, in the same way, as
# `resolve_canonical_rank1_lineage` selects the strategy the chapter goes on to describe.

# %% tags=["results"]
set_global_seeds(SEED)
universe_symbols = yaml.safe_load(
    (get_case_study_dir(CASE_STUDY_ID) / "config" / "setup.yaml").read_text()
)["universe"]["symbols"]
n_assets = len(universe_symbols)
if SPLIT != "validation":
    raise ValueError("cost sensitivity uses validation backtests")
if FORCE_REBACKTEST:
    raise ValueError("identical complete backtests are reused by identity")
if not RUN_SWEEP:
    raise ValueError("set RUN_SWEEP=True to execute the visible cost request")

study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
# The execution tier decides which registry namespace this run reads and writes;
# the reduction knobs decide only how much of it is covered. Inferring the tier
# from the knobs conflated the two, so any reduced run went looking for preview
# predictions - and a reduced run over a canonical upstream, which is what the
# test suite exercises, then resolved no rows at all.
include_preview = EXECUTION_TIER == "preview"

# The tier decides the namespace, so a canonical run may legitimately be narrowed -
# but a narrowed run declares a different set of members than the canonical
# population does, and a population is immutable once written. Such a run must
# publish under its own name rather than register a partial snapshot of the cost sweep
# under the canonical one.
if (
    (TOP_K or TOP_N_PREDICTIONS is not None or MAX_COST_POINTS or LABEL)
    and not include_preview
    and not POPULATION_NAME
):
    raise ValueError(
        "this run narrows the cost sweep, so it cannot publish the canonical "
        "population; pass POPULATION_NAME to give it its own"
    )
catalog = study.predictions.table(include_preview=include_preview).filter(
    (pl.col("identity_status") == "current")
    & (pl.col("split") == SPLIT)
    & pl.col("complete")
    & (pl.col("execution_tier") == ("preview" if include_preview else "canonical"))
)
# `identity_status` is the schema version a row was written under, not a statement about which
# generation its producer publishes. A model notebook that refits leaves the generation it
# replaced in the registry, complete and current, so this filter alone would carry a retired
# prediction set into the sweep. `superseded_members` reads the lineage instead - see
# `13_backtest`, which drops the same set before it freezes the baseline population.
# `SUPERSEDES_COST_BACKTESTS` names the snapshot this run replaces under the name it publishes,
# offered through `population_supersedes` on the same rule. It is empty until that name has a
# first generation; after that, an upstream refit changes this population's member list and
# the registry refuses the write without it. `13_backtest` states the reasoning once.
retired = superseded_members(study, member_kind="prediction")
if retired:
    catalog = catalog.filter(~pl.col("prediction_hash").is_in(list(retired)))
if LABEL:
    catalog = catalog.filter(pl.col("label") == LABEL)
if TOP_N_PREDICTIONS is not None:
    catalog = catalog.sort("label", "family", "config_name", "checkpoint_value").head(
        TOP_N_PREDICTIONS
    )
if catalog.is_empty():
    raise RuntimeError("cost sensitivity resolved no complete prediction rows")


def _open_backtests(population: OfficialPopulation) -> list[BacktestResult]:
    population.require_complete()
    opened = [Result.open(study, value) for value in population.members]
    if any(not isinstance(result, BacktestResult) for result in opened):
        raise TypeError(f"population {population.name!r} contains a non-backtest result")
    return [result for result in opened if isinstance(result, BacktestResult)]


def _label(result: BacktestResult) -> str:
    return str(result.lineage()["training_spec"]["label"])


def _registered_preview_allocations() -> pl.DataFrame:
    """The allocation backtests an upstream preview registered.

    One predicate, read once. Stating it twice - once to decide which labels are covered
    and once to pick a label's leader - makes the two required to agree forever: tighten
    one and the other admits a label whose backtests it then rejects, which is exactly the
    "no preview allocation backtests are registered" failure this exists to prevent.
    """
    return study.backtests.table(include_preview=True).filter(
        (pl.col("stage") == "allocation")
        & (pl.col("execution_tier") == "preview")
        & (pl.col("identity_status") == "current")
        & pl.col("complete")
    )


def _preview_leader(rows: pl.DataFrame, registered_allocations: pl.DataFrame) -> BacktestResult:
    """One allocation backtest for this label, read from what 14 registered.

    Rebuilding the identity restated two of the upstream run's choices - its `top_k` and
    which allocator came first - from this notebook's own defaults. A preview reduces
    each notebook independently, so those are guesses about another run's parameters,
    and a guess that is wrong looks for a hash nothing wrote instead of reporting the
    disagreement.
    """
    registered = registered_allocations.filter(
        pl.col("prediction_hash").is_in(rows.get_column("prediction_hash").implode())
    )
    if registered.is_empty():
        raise RuntimeError(
            "no preview allocation backtests are registered for this prediction "
            "catalog; run 14_portfolio_management at the same reduction first"
        )
    row = registered.sort("family", "config_name", "checkpoint_value", "backtest_hash").row(
        0, named=True
    )
    result = Result.open(study, row["backtest_hash"], include_preview=True)
    if not isinstance(result, BacktestResult) or not result.complete:
        raise RuntimeError("the deterministic preview allocation result is not complete")
    return result


selected_by_label: dict[str, BacktestResult] = {}
candidate_sets: dict[str, CandidateSet] = {}
if include_preview:
    # The labels come from what the upstream preview registered, the same rule the
    # canonical branch below follows. Enumerating this notebook's own catalog asks
    # _preview_leader for labels 14_portfolio_management was configured not to allocate,
    # and it raises "no preview allocation backtests are registered" for each - reporting
    # a reduction the run was told to make as a missing upstream.
    registered_allocations = _registered_preview_allocations()
    covered = catalog.filter(
        pl.col("prediction_hash").is_in(
            registered_allocations.get_column("prediction_hash").implode()
        )
    )
    if covered.is_empty():
        raise RuntimeError(
            "no preview allocation backtests cover this prediction catalog; "
            "run 14_portfolio_management at the same reduction first"
        )
    for label in sorted(covered.get_column("label").unique()):
        selected_by_label[label] = _preview_leader(
            covered.filter(pl.col("label") == label), registered_allocations
        )
else:
    baselines = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "equal-weight-baselines", scope=POPULATION_NAME),
        )
    )
    allocations = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "allocation-backtests", scope=POPULATION_NAME),
        )
    )
    risk_overlays = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "risk-overlay-backtests", scope=POPULATION_NAME),
        )
    )
    # The labels come from the upstream populations this run resolved, not from this
    # notebook's own catalog. A narrowed upstream covers fewer labels than the catalog
    # holds, and rebuilding the list locally reproduces the upstream narrowing by
    # convention. Unscoped, the run publishes canonical names and the two must agree.
    upstream = [*baselines, *allocations, *risk_overlays]
    upstream_labels = sorted({_label(result) for result in upstream})
    if not upstream_labels:
        raise RuntimeError("the upstream populations carry no labels")
    if not POPULATION_NAME and upstream_labels != sorted(catalog.get_column("label").unique()):
        raise RuntimeError(
            "the canonical upstream populations do not cover every label in the catalog: "
            f"upstream {upstream_labels}, "
            f"catalog {sorted(catalog.get_column('label').unique())}"
        )
    # Eligibility and order both come from `selectable_validation_candidates`, which is the
    # function `resolve_solvent_carrier` ranks. Re-deriving them here is what put the cost curve
    # on the wrong strategy: the three populations above are read whole, and the retired-prediction
    # filter a few cells up is applied to the *catalog* and never to *them*. `56070f34dff1` is a
    # published risk-overlay backtest whose prediction `9eb5f506a0ee` was superseded by a refit, so
    # it survived here, won on raw Sharpe, and eleven cost points were swept over a strategy the
    # case study does not report - while `19_strategy_analysis`, which asks the resolver, reported
    # `747e7e47abaa`. Nothing raised, because a cost sweep over the wrong parent is a valid sweep.
    #
    # The ordering matters too, not only the eligibility. Where a conformal candidate is in the
    # field the resolver re-ranks every member on the timestamps they all price, because a
    # calibration that abstains through its warm-up books those decisions as zero and is otherwise
    # compared against allocators measured over a longer span. `best_validation_sharpe()` sorts on
    # the stored number and does neither.
    _eligible_order = {
        row["backtest_hash"]: position
        for position, row in enumerate(
            selectable_validation_candidates(CASE_STUDY_ID, labels=[LABEL] if LABEL else None)
        )
    }
    for label in upstream_labels:
        members = [result for result in upstream if _label(result) == label]
        eligible = [result for result in members if result.hash in _eligible_order]
        if not eligible:
            raise RuntimeError(
                f"none of the {len(members)} upstream backtests for {label} is selectable: "
                "every one is retired on the backtest or the prediction side, or belongs to no "
                "population its producer publishes. Re-run the validation stages rather than "
                "sweeping costs over a strategy nothing reports."
            )
        _set_name = research_name(
            CASE_STUDY_ID, f"{label}:pre-cost-strategies", scope=POPULATION_NAME
        )
        # The frozen set records the field the selection actually saw, so it holds the
        # selectable members and not every row the three populations list.
        candidates = CandidateSet.create(
            study,
            name=_set_name,
            members=eligible,
            supersedes=candidate_set_supersedes(
                study, name=_set_name, declared=SUPERSEDES_CANDIDATE_SETS.get(_set_name)
            ),
        )
        candidate_sets[label] = candidates
        leader = min(eligible, key=lambda result: _eligible_order[result.hash])
        if not isinstance(leader, BacktestResult):
            raise TypeError("strategy selection did not return a backtest")
        selected_by_label[label] = leader


def _overlay_gap(label: str) -> dict[str, object]:
    """What the risk controls were worth on *label*, in validation Sharpe.

    The parent is chosen across all three stages, so "the overlay won" and "the overlay
    helped" are the same statement only when an un-overlaid configuration was in the running.
    Both sides are reported: the best risk overlay, the best signal-or-allocation strategy it
    was measured against, and the difference. A negative difference is the finding that the
    controls cost more than they saved on this label, and the selected parent is then
    un-overlaid.
    """
    members = candidate_sets[label].members
    rows = study.backtests.table().filter(
        pl.col("backtest_hash").is_in(list(members))
        & (pl.col("split") == "validation")
        & pl.col("sharpe").is_not_null()
    )
    overlaid = rows.filter(pl.col("stage") == "risk_overlay")
    un_overlaid = rows.filter(pl.col("stage").is_in(["signal", "allocation"]))
    best_overlaid = overlaid.get_column("sharpe").max() if overlaid.height else None
    best_un_overlaid = un_overlaid.get_column("sharpe").max() if un_overlaid.height else None
    gap = (
        best_overlaid - best_un_overlaid
        if best_overlaid is not None and best_un_overlaid is not None
        else None
    )
    return {
        "best_overlaid_sharpe": best_overlaid,
        "best_un_overlaid_sharpe": best_un_overlaid,
        "overlay_gap": gap,
    }


pl.DataFrame(
    [
        {
            "label": label,
            "backtest_hash": result.hash,
            "prediction_hash": result.registry_record()["prediction_hash"],
            "stage": result.registry_record()["stage"],
            **(_overlay_gap(label) if label in candidate_sets else {}),
        }
        for label, result in selected_by_label.items()
    ]
)

# %% [markdown]
# ## Plan and freeze exact cost siblings
#
# The configured cost grid is expressed as total basis points per traded leg. Commission and
# slippage each receive half. The identity audit removes only the cost fields and the chapter label;
# every remaining field must match the selected validation strategy. Production freezes the full
# sensitivity set before the first backtest is written.
#
# The even split between commission and slippage is a modelling convention, not a measurement.
# Nothing in the data says the two halves of the friction are equal; the split exists so that a
# single configured rate can populate two fields the backtest engine charges separately. Read the
# total, not the halves.
#
# What the curve measures, once it exists, is turnover. Cost enters the return series through
# `|delta w|` at each rebalance, so a strategy's sensitivity to the assumed rate is set by how
# much of the book it moves and how often, not by how good its predictions are. Two strategies
# with the same gross Sharpe can have breakevens that differ by an order of magnitude, and the
# whole reason to plot a curve rather than report one number is that the difference is invisible
# at any single rate. The quantity to read off is where the curve crosses zero and how far that
# sits from the quoted band above - a strategy that survives to 40 basis points on pairs that
# trade at 3 has room; one that dies at 4 is reporting an edge that is really a spread.

# %% tags=["results"]
cost_grid = get_cost_grid_bps(CASE_STUDY_ID)
if MAX_COST_POINTS:
    cost_grid = cost_grid[:MAX_COST_POINTS]
if not cost_grid:
    raise RuntimeError("the cost grid is empty")


def _catalog_row(result: BacktestResult) -> pl.DataFrame:
    prediction_hash = result.registry_record()["prediction_hash"]
    row = catalog.filter(pl.col("prediction_hash") == prediction_hash)
    if row.height != 1:
        raise RuntimeError(f"prediction {prediction_hash} resolved to {row.height} catalog rows")
    return row


def _strategy_arguments(result: BacktestResult) -> dict[str, Any]:
    strategy = result.spec()["strategy"]
    return {
        "signal": deepcopy(strategy["signal"]),
        "allocation": deepcopy(strategy.get("allocation")),
        "risk": deepcopy(strategy.get("risk")),
        "execution_mode": strategy.get("rebalance", {}).get("mode"),
    }


def _non_cost_projection(spec: dict[str, Any]) -> dict[str, Any]:
    projected = deepcopy(spec)
    projected.pop("chapter", None)
    projected.pop("_runtime_backtest_config", None)
    config = projected.get("backtest_config", {})
    config.pop("commission", None)
    config.pop("slippage", None)
    metadata = config.get("metadata")
    if isinstance(metadata, dict):
        metadata.pop("chapter", None)
        # An absolute filesystem path, and `case_studies/utils/registry/specs.py` already
        # excludes it from the identity hash for that reason. Comparing it here made the
        # check fail on where the notebook was run from rather than on what it produced:
        # a sibling written in one worktree never matches a parent registered in another,
        # and the message says a strategy field moved when none did.
        metadata.pop("preset_path", None)
    # The parent was serialized by whatever engine version registered it and the sibling by the
    # installed one, so a field the engine has since added to `BacktestConfig` is present on one
    # side and absent on the other while both describe the same strategy. `ml4t-backtest` 0.1.3 to
    # 0.1.6 added `account.lock_notional_update_mode` and `position_sizing.share_rounding`, which
    # failed every fx_pairs and crypto_perps_funding parent registered before 2026-09-12 with a
    # message saying a strategy field moved. Round-tripping both sides through the installed schema
    # states the comparison in one vocabulary, so the check answers what this notebook built rather
    # than which engine wrote the row it is compared against - and it covers the next added field
    # without naming it. `ensure_backtest_spec` deliberately does NOT round-trip, because there the
    # result is hashed and dropped unknown keys would move an identity; here it is compared and
    # discarded. Metadata is merged back over the serialized view for the same reason it is there:
    # the dataclass pins a schema and drops keys it does not know.
    if EngineBacktestConfig is not None and config:
        original_metadata = dict(metadata) if isinstance(metadata, dict) else {}
        rebuilt = EngineBacktestConfig.from_dict(config).to_dict()
        rebuilt.pop("commission", None)
        rebuilt.pop("slippage", None)
        rebuilt_metadata = dict(rebuilt.get("metadata") or {})
        rebuilt_metadata.update(original_metadata)
        rebuilt["metadata"] = rebuilt_metadata
        projected["backtest_config"] = rebuilt
    return projected


cost_jobs = []
for label, selected in selected_by_label.items():
    arguments = _strategy_arguments(selected)
    for total_bps in cost_grid:
        costs = {
            "commission_bps": total_bps / 2.0,
            "slippage_bps": total_bps / 2.0,
        }
        plan = plan_backtests(
            study,
            predictions=_catalog_row(selected),
            signal=arguments["signal"],
            allocation=arguments["allocation"],
            risk=arguments["risk"],
            costs=costs,
            chapter="ch18",
            execution_mode=arguments["execution_mode"],
        )
        if len(plan.members) != 1:
            raise RuntimeError("a cost plan must contain exactly one backtest")
        cost_jobs.append(
            {
                "label": label,
                "selected": selected,
                "arguments": arguments,
                "total_bps": total_bps,
                "costs": costs,
                "backtest_hash": plan.expected_hashes[0],
            }
        )

planned_hashes = [job["backtest_hash"] for job in cost_jobs]
if len(planned_hashes) != len(set(planned_hashes)):
    raise RuntimeError("two planned cost requests collapse to the same identity")

cost_population = None
if not include_preview:
    costs_name = research_name(CASE_STUDY_ID, "cost-sensitivity-backtests", scope=POPULATION_NAME)
    cost_population = OfficialPopulation.create(
        study,
        name=costs_name,
        member_kind="backtest",
        members=planned_hashes,
        supersedes=population_supersedes(
            study, name=costs_name, declared=SUPERSEDES_COST_BACKTESTS
        ),
    )
    print(f"Frozen expected cost population: {cost_population.hash}")

# %% [markdown]
# ## Execute the frozen cost grid, and validate its membership without making it selectable
#
# The population is validated in the cell that fills it: the expected set was written down before
# the first member ran, and `require_complete` is what turns that declaration into a published
# result. Publishing it does not make it selectable - a cost sensitivity is a curve through a
# parameter the strategy does not choose, and later selection reads the allocation population.
#
# "Registered but not selectable" is a distinction worth being concrete about, because both parts
# are deliberate. These rows are written to the registry, complete and current, exactly like every
# other backtest: the curve is a published result that a reader can look up and re-derive. What
# makes them unselectable is that the downstream stages read named populations rather than
# querying the registry for whatever is complete, so a cost sibling is never a member of a set
# anything ranks. The separation lives in which population a stage reads, not in a flag on the
# row - which is why a stage that queried the registry directly would silently acquire eleven
# copies of one strategy, each at a different assumed rate, and would rank them.

# %% tags=["results"]
# A sweep that recomputes everything and a sweep that recomputes nothing print the same summary
# unless the two are counted apart. `run_backtests` serves an identity that is already registered
# and complete instead of running it again, which is what makes a re-run affordable and what makes
# a bare member count say nothing about whether this run did any work.
#
# The runner already knows which it did and says so per member in `execution.diagnostics`, as
# `status` "reused" or "completed". Comparing against the registered hashes instead would be
# wrong in both directions: a registered-but-partial backtest is in that set, gets recomputed and
# would report as reused, and a preview re-run reads a table that excludes preview rows by default
# and would report every reused member as computed.
run_status: list[str] = []
cost_results: list[BacktestResult] = []
cost_rows = []
for job in cost_jobs:
    selected = job["selected"]
    arguments = job["arguments"]
    execution = run_backtests(
        study,
        predictions=_catalog_row(selected),
        signal=arguments["signal"],
        allocation=arguments["allocation"],
        risk=arguments["risk"],
        costs=job["costs"],
        chapter="ch18",
        execution_mode=arguments["execution_mode"],
    )
    if len(execution.results) != 1:
        raise RuntimeError("a cost request must produce exactly one backtest")
    result = execution.results[0]
    if result.hash != job["backtest_hash"]:
        raise RuntimeError("a completed cost identity differs from the frozen plan")
    if _non_cost_projection(result.spec()) != _non_cost_projection(selected.spec()):
        raise RuntimeError("a cost sibling changed a non-cost strategy field")
    if result.registry_record()["stage"] != "cost_sensitivity" or not result.complete:
        raise RuntimeError("a cost result is incomplete or misclassified")
    cost_results.append(result)
    run_status.extend(entry["status"] for entry in execution.diagnostics)
    cost_rows.append(
        {
            "label": job["label"],
            "total_cost_bps": job["total_bps"],
            "backtest_hash": result.hash,
            "prediction_hash": result.registry_record()["prediction_hash"],
        }
    )

served = run_status.count("reused")
print(
    f"Cost siblings: {reuse_disclosure(len(cost_results) - served, served)}, "
    f"{len(cost_results)} in the population"
)

if not include_preview:
    if cost_population is None:
        raise RuntimeError("the canonical cost population was not frozen before execution")
    cost_population.require_complete()
    print(f"Official cost-sensitivity population: {cost_population.hash}")
else:
    print("Preview cost curves remain outside official populations and candidate sets.")

pl.DataFrame(cost_rows).sort("label", "total_cost_bps")

# %% [markdown]
# ## Key takeaways
#
# - Each label has one validation-selected parent strategy, chosen across all three upstream
#   stages so the curve prices the strategy the chapter reports rather than a sibling of it.
# - Cost siblings preserve every non-cost identity field.
# - Cost sensitivity is frozen for completeness but excluded from later selection. A strategy that
#   could compete on its cost assumption would win by assuming costs away.
# - Basis points are the FX regime because spreads are quoted as a fraction of the rate. The
#   declared components are a taxonomy; the backtest charges one aggregate rate per traded leg.
# - The curve is a turnover measurement. Read where it crosses zero and compare that to the
#   quoted band, not the Sharpe at any single rate.
#
# The breakeven this produces is still a validation-period number, and turnover is not stable
# across regimes: a strategy that trades more in volatile periods pays more exactly when spreads
# are widest, and a curve computed at a constant rate cannot show that. The curve bounds the
# question rather than settling it.

```

Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT

Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.