Pular para o conteúdo
Todos os documentos da biblioteca

Sensibilidade de estratégia FX aos custos de transação

Código Machine Learning for Trading

Resumo

Este notebook mede como uma estratégia FX selecionada na validação responde a mudanças nos custos de transação proporcionais. Para cada rótulo de previsão, seleciona uma estratégia de referência a partir dos resultados de sinal, alocação e camada de risco e, em seguida, executa novamente essa estratégia em uma grade de custos, mantendo os demais campos de identidade. A variação de custos é uma análise de perturbação: descreve a robustez de uma estratégia definida e não participa de seleções posteriores.

O documento explica por que pontos-base são adequados para spreads de FX, que são proporcionais à taxa de câmbio, e por que o backtest aplica um custo agregado por perna negociada, embora a configuração liste separadamente os componentes de spread e swap. A curva é interpretada por meio do giro da carteira e de seu cruzamento com zero em relação aos níveis de spread cotados. Esses resultados são do período de validação, e as hipóteses de taxa constante não representam spreads que aumentam junto com a volatilidade nem mudanças no giro entre regimes. O trecho apresenta um método e suas ressalvas, mas nenhum resultado realizado da estratégia.

Ideias principais

  • Escolha a estratégia de referência entre os candidatos de validação nas etapas de sinal, alocação e risco antes de testar os custos.
  • Mantenha as variações de custo fora das seleções posteriores para evitar favorecer hipóteses de fricção artificialmente baixa.
  • Represente os custos de FX de forma proporcional, pois os spreads são expressos em relação às taxas de câmbio.
  • Interprete a sensibilidade a custos como uma relação entre giro, desempenho líquido e taxa de custo presumida.
  • Uma curva de custos constante no período de validação não captura mudanças de regime no giro nem o aumento dos spreads.

Tags

Texto completo
# 16_costs.py


```py
# ---
# jupyter:
#   jupytext:
#     cell_metadata_filter: tags,-all
#     text_representation:
#       extension: .py
#       format_name: percent
#       format_version: '1.3'
#       jupytext_version: 1.19.3
#   kernelspec:
#     display_name: Python 3 (ipykernel)
#     language: python
#     name: python3
# ---

# %% [markdown]
# # Transaction-Cost Sensitivity - FX Pairs
#
# This notebook selects one validation strategy per label from the equal-weight, allocation and
# risk-overlay populations, then changes only its percentage transaction costs. The sensitivity
# grid does not participate in later model or strategy selection.
#
# Costs are swept last because a cost curve is only informative about the strategy that would
# actually be traded, and that strategy is not settled until the risk controls have been measured.
# Sweeping before the overlay charges the grid against a configuration the next notebook may
# discard.
#
# This is a perturbation analysis, not a choice. It asks what happens to a settled strategy if the
# cost model is wrong, which is a different question from which strategy to trade. The two must
# stay separate for a mechanical reason: a strategy allowed to compete on its own cost assumption
# would win by having costs assumed away, and the ranking would report the most optimistic
# assumption rather than the best strategy.
#
# **Why basis points here.** FX spreads are quoted in pips, a fraction of the rate itself, so a
# proportional charge is how the friction is actually expressed - there is no share or contract
# to bill per unit. Case studies where nominal prices are stable, or where spreads are measured
# from quote data, use per-share instead; applying a flat per-unit charge to a rate would assume
# spread scales with the level, which it does not. The configured grid runs from 0 to 50 basis
# points per traded leg, against a real quoted band of roughly 1 to 3 for major pairs and 3 to 8
# for crosses, so most of the grid sits deliberately past anything plausible.
#
# `config/setup.yaml` names the cost components as spread and swap points. That list is a
# taxonomy, read by the Chapter 18 teaching notebooks to describe what the friction consists of;
# the backtest charges a single aggregate rate per traded leg. So this curve perturbs the
# aggregate, and the overnight financing cost of carrying a position is described rather than
# separately priced.
#
# **Learning objectives**
#
# - Select from an immutable validation candidate set by backtest Sharpe.
# - Preserve model, checkpoint, signal, allocation, and execution identities across a cost curve.
# - Keep cost sensitivity outside the official selection cohort.
# - Read a cost curve as a statement about turnover.
#
# **Book reference**: Chapter 18
#
# **Prerequisite**: `15_risk_management`.

# %%
"""Run one cost-sensitivity curve per FX prediction label."""

from copy import deepcopy
from typing import Any

import polars as pl
import yaml

from case_studies.research import (
    BacktestResult,
    CandidateSet,
    OfficialPopulation,
    Result,
    candidate_set_supersedes,
    open_study,
    plan_backtests,
    population_supersedes,
    research_name,
    reuse_disclosure,
    run_backtests,
    superseded_members,
)
from case_studies.utils.backtest_presets import EngineBacktestConfig
from case_studies.utils.strategy_analysis import selectable_validation_candidates
from case_studies.utils.sweep_config import (
    get_cost_grid_bps,
)
from utils.paths import get_case_study_dir
from utils.reproducibility import set_global_seeds

# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
SPLIT = "validation"
TOP_K = 0
TOP_N_PREDICTIONS = None
MAX_COST_POINTS = 0
SEED = 42
RUN_SWEEP = True
FORCE_REBACKTEST = False
POPULATION_NAME = ""
SUPERSEDES_COST_BACKTESTS: str = "live"
# The same rule the populations follow: a candidate set is immutable under its name, so a rebuilt
# upstream generation must name the set it replaces. Keyed by the full set name, which is what the
# refusal prints. `15_risk_management` states the reasoning once.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "fx_pairs:fwd_ret_1d:pre-cost-strategies": "live",
    "fx_pairs:fwd_ret_5d:pre-cost-strategies": "live",
    "fx_pairs:fwd_ret_21d:pre-cost-strategies": "live",
}

# %% [markdown]
# ## Select one strategy for each label
#
# Production selection considers the signal, allocation and risk-overlay populations together -
# the same three stages the canonical validation rank-1 is selected over in
# `case_studies/utils/strategy_analysis.py:resolve_canonical_rank1_lineage`. All three are
# candidates because each stage is an alternative to the one before it rather than an improvement
# on it by construction. Naming only the overlays would charge the cost grid against a risk
# control even where every control measured worse than leaving the position rule alone; naming
# only signal and allocation would sweep a strategy the risk notebook has already improved on.
# Which stage the parent came from is printed below rather than assumed, and the gap between the
# best overlay and the best un-overlaid configuration is printed with it: a negative gap is the
# measurement that the controls did not help on this label.
#
# Cost variants are descendants of that choice and cannot improve their own chance of selection.
# Preview mode uses a deterministic allocation request from the reduced catalog and remains
# outside candidate sets.
#
# Getting the parent wrong has a quiet failure mode. The cost curve would be computed correctly,
# the population would freeze and validate, and every number would be right - about a strategy
# the chapter does not report. Nothing raises, because a cost sweep over the wrong parent is a
# perfectly valid sweep. That is why the stage the parent came from is printed rather than
# assumed, and why the selection here is made over the same three stages, in the same way, as
# `resolve_canonical_rank1_lineage` selects the strategy the chapter goes on to describe.

# %% tags=["results"]
set_global_seeds(SEED)
universe_symbols = yaml.safe_load(
    (get_case_study_dir(CASE_STUDY_ID) / "config" / "setup.yaml").read_text()
)["universe"]["symbols"]
n_assets = len(universe_symbols)
if SPLIT != "validation":
    raise ValueError("cost sensitivity uses validation backtests")
if FORCE_REBACKTEST:
    raise ValueError("identical complete backtests are reused by identity")
if not RUN_SWEEP:
    raise ValueError("set RUN_SWEEP=True to execute the visible cost request")

study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
# The execution tier decides which registry namespace this run reads and writes;
# the reduction knobs decide only how much of it is covered. Inferring the tier
# from the knobs conflated the two, so any reduced run went looking for preview
# predictions - and a reduced run over a canonical upstream, which is what the
# test suite exercises, then resolved no rows at all.
include_preview = EXECUTION_TIER == "preview"

# The tier decides the namespace, so a canonical run may legitimately be narrowed -
# but a narrowed run declares a different set of members than the canonical
# population does, and a population is immutable once written. Such a run must
# publish under its own name rather than register a partial snapshot of the cost sweep
# under the canonical one.
if (
    (TOP_K or TOP_N_PREDICTIONS is not None or MAX_COST_POINTS or LABEL)
    and not include_preview
    and not POPULATION_NAME
):
    raise ValueError(
        "this run narrows the cost sweep, so it cannot publish the canonical "
        "population; pass POPULATION_NAME to give it its own"
    )
catalog = study.predictions.table(include_preview=include_preview).filter(
    (pl.col("identity_status") == "current")
    & (pl.col("split") == SPLIT)
    & pl.col("complete")
    & (pl.col("execution_tier") == ("preview" if include_preview else "canonical"))
)
# `identity_status` is the schema version a row was written under, not a statement about which
# generation its producer publishes. A model notebook that refits leaves the generation it
# replaced in the registry, complete and current, so this filter alone would carry a retired
# prediction set into the sweep. `superseded_members` reads the lineage instead - see
# `13_backtest`, which drops the same set before it freezes the baseline population.
# `SUPERSEDES_COST_BACKTESTS` names the snapshot this run replaces under the name it publishes,
# offered through `population_supersedes` on the same rule. It is empty until that name has a
# first generation; after that, an upstream refit changes this population's member list and
# the registry refuses the write without it. `13_backtest` states the reasoning once.
retired = superseded_members(study, member_kind="prediction")
if retired:
    catalog = catalog.filter(~pl.col("prediction_hash").is_in(list(retired)))
if LABEL:
    catalog = catalog.filter(pl.col("label") == LABEL)
if TOP_N_PREDICTIONS is not None:
    catalog = catalog.sort("label", "family", "config_name", "checkpoint_value").head(
        TOP_N_PREDICTIONS
    )
if catalog.is_empty():
    raise RuntimeError("cost sensitivity resolved no complete prediction rows")


def _open_backtests(population: OfficialPopulation) -> list[BacktestResult]:
    population.require_complete()
    opened = [Result.open(study, value) for value in population.members]
    if any(not isinstance(result, BacktestResult) for result in opened):
        raise TypeError(f"population {population.name!r} contains a non-backtest result")
    return [result for result in opened if isinstance(result, BacktestResult)]


def _label(result: BacktestResult) -> str:
    return str(result.lineage()["training_spec"]["label"])


def _registered_preview_allocations() -> pl.DataFrame:
    """The allocation backtests an upstream preview registered.

    One predicate, read once. Stating it twice - once to decide which labels are covered
    and once to pick a label's leader - makes the two required to agree forever: tighten
    one and the other admits a label whose backtests it then rejects, which is exactly the
    "no preview allocation backtests are registered" failure this exists to prevent.
    """
    return study.backtests.table(include_preview=True).filter(
        (pl.col("stage") == "allocation")
        & (pl.col("execution_tier") == "preview")
        & (pl.col("identity_status") == "current")
        & pl.col("complete")
    )


def _preview_leader(rows: pl.DataFrame, registered_allocations: pl.DataFrame) -> BacktestResult:
    """One allocation backtest for this label, read from what 14 registered.

    Rebuilding the identity restated two of the upstream run's choices - its `top_k` and
    which allocator came first - from this notebook's own defaults. A preview reduces
    each notebook independently, so those are guesses about another run's parameters,
    and a guess that is wrong looks for a hash nothing wrote instead of reporting the
    disagreement.
    """
    registered = registered_allocations.filter(
        pl.col("prediction_hash").is_in(rows.get_column("prediction_hash").implode())
    )
    if registered.is_empty():
        raise RuntimeError(
            "no preview allocation backtests are registered for this prediction "
            "catalog; run 14_portfolio_management at the same reduction first"
        )
    row = registered.sort("family", "config_name", "checkpoint_value", "backtest_hash").row(
        0, named=True
    )
    result = Result.open(study, row["backtest_hash"], include_preview=True)
    if not isinstance(result, BacktestResult) or not result.complete:
        raise RuntimeError("the deterministic preview allocation result is not complete")
    return result


selected_by_label: dict[str, BacktestResult] = {}
candidate_sets: dict[str, CandidateSet] = {}
if include_preview:
    # The labels come from what the upstream preview registered, the same rule the
    # canonical branch below follows. Enumerating this notebook's own catalog asks
    # _preview_leader for labels 14_portfolio_management was configured not to allocate,
    # and it raises "no preview allocation backtests are registered" for each - reporting
    # a reduction the run was told to make as a missing upstream.
    registered_allocations = _registered_preview_allocations()
    covered = catalog.filter(
        pl.col("prediction_hash").is_in(
            registered_allocations.get_column("prediction_hash").implode()
        )
    )
    if covered.is_empty():
        raise RuntimeError(
            "no preview allocation backtests cover this prediction catalog; "
            "run 14_portfolio_management at the same reduction first"
        )
    for label in sorted(covered.get_column("label").unique()):
        selected_by_label[label] = _preview_leader(
            covered.filter(pl.col("label") == label), registered_allocations
        )
else:
    baselines = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "equal-weight-baselines", scope=POPULATION_NAME),
        )
    )
    allocations = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "allocation-backtests", scope=POPULATION_NAME),
        )
    )
    risk_overlays = _open_backtests(
        OfficialPopulation.one(
            study,
            name=research_name(CASE_STUDY_ID, "risk-overlay-backtests", scope=POPULATION_NAME),
        )
    )
    # The labels come from the upstream populations this run resolved, not from this
    # notebook's own catalog. A narrowed upstream covers fewer labels than the catalog
    # holds, and rebuilding the list locally reproduces the upstream narrowing by
    # convention. Unscoped, the run publishes canonical names and the two must agree.
    upstream = [*baselines, *allocations, *risk_overlays]
    upstream_labels = sorted({_label(result) for result in upstream})
    if not upstream_labels:
        raise RuntimeError("the upstream populations carry no labels")
    if not POPULATION_NAME and upstream_labels != sorted(catalog.get_column("label").unique()):
        raise RuntimeError(
            "the canonical upstream populations do not cover every label in the catalog: "
            f"upstream {upstream_labels}, "
            f"catalog {sorted(catalog.get_column('label').unique())}"
        )
    # Eligibility and order both come from `selectable_validation_candidates`, which is the
    # function `resolve_solvent_carrier` ranks. Re-deriving them here is what put the cost curve
    # on the wrong strategy: the three populations above are read whole, and the retired-prediction
    # filter a few cells up is applied to the *catalog* and never to *them*. `56070f34dff1` is a
    # published risk-overlay backtest whose prediction `9eb5f506a0ee` was superseded by a refit, so
    # it survived here, won on raw Sharpe, and eleven cost points were swept over a strategy the
    # case study does not report - while `19_strategy_analysis`, which asks the resolver, reported
    # `747e7e47abaa`. Nothing raised, because a cost sweep over the wrong parent is a valid sweep.
    #
    # The ordering matters too, not only the eligibility. Where a conformal candidate is in the
    # field the resolver re-ranks every member on the timestamps they all price, because a
    # calibration that abstains through its warm-up books those decisions as zero and is otherwise
    # compared against allocators measured over a longer span. `best_validation_sharpe()` sorts on
    # the stored number and does neither.
    _eligible_order = {
        row["backtest_hash"]: position
        for position, row in enumerate(
            selectable_validation_candidates(CASE_STUDY_ID, labels=[LABEL] if LABEL else None)
        )
    }
    for label in upstream_labels:
        members = [result for result in upstream if _label(result) == label]
        eligible = [result for result in members if result.hash in _eligible_order]
        if not eligible:
            raise RuntimeError(
                f"none of the {len(members)} upstream backtests for {label} is selectable: "
                "every one is retired on the backtest or the prediction side, or belongs to no "
                "population its producer publishes. Re-run the validation stages rather than "
                "sweeping costs over a strategy nothing reports."
            )
        _set_name = research_name(
            CASE_STUDY_ID, f"{label}:pre-cost-strategies", scope=POPULATION_NAME
        )
        # The frozen set records the field the selection actually saw, so it holds the
        # selectable members and not every row the three populations list.
        candidates = CandidateSet.create(
            study,
            name=_set_name,
            members=eligible,
            supersedes=candidate_set_supersedes(
                study, name=_set_name, declared=SUPERSEDES_CANDIDATE_SETS.get(_set_name)
            ),
        )
        candidate_sets[label] = candidates
        leader = min(eligible, key=lambda result: _eligible_order[result.hash])
        if not isinstance(leader, BacktestResult):
            raise TypeError("strategy selection did not return a backtest")
        selected_by_label[label] = leader


def _overlay_gap(label: str) -> dict[str, object]:
    """What the risk controls were worth on *label*, in validation Sharpe.

    The parent is chosen across all three stages, so "the overlay won" and "the overlay
    helped" are the same statement only when an un-overlaid configuration was in the running.
    Both sides are reported: the best risk overlay, the best signal-or-allocation strategy it
    was measured against, and the difference. A negative difference is the finding that the
    controls cost more than they saved on this label, and the selected parent is then
    un-overlaid.
    """
    members = candidate_sets[label].members
    rows = study.backtests.table().filter(
        pl.col("backtest_hash").is_in(list(members))
        & (pl.col("split") == "validation")
        & pl.col("sharpe").is_not_null()
    )
    overlaid = rows.filter(pl.col("stage") == "risk_overlay")
    un_overlaid = rows.filter(pl.col("stage").is_in(["signal", "allocation"]))
    best_overlaid = overlaid.get_column("sharpe").max() if overlaid.height else None
    best_un_overlaid = un_overlaid.get_column("sharpe").max() if un_overlaid.height else None
    gap = (
        best_overlaid - best_un_overlaid
        if best_overlaid is not None and best_un_overlaid is not None
        else None
    )
    return {
        "best_overlaid_sharpe": best_overlaid,
        "best_un_overlaid_sharpe": best_un_overlaid,
        "overlay_gap": gap,
    }


pl.DataFrame(
    [
        {
            "label": label,
            "backtest_hash": result.hash,
            "prediction_hash": result.registry_record()["prediction_hash"],
            "stage": result.registry_record()["stage"],
            **(_overlay_gap(label) if label in candidate_sets else {}),
        }
        for label, result in selected_by_label.items()
    ]
)

# %% [markdown]
# ## Plan and freeze exact cost siblings
#
# The configured cost grid is expressed as total basis points per traded leg. Commission and
# slippage each receive half. The identity audit removes only the cost fields and the chapter label;
# every remaining field must match the selected validation strategy. Production freezes the full
# sensitivity set before the first backtest is written.
#
# The even split between commission and slippage is a modelling convention, not a measurement.
# Nothing in the data says the two halves of the friction are equal; the split exists so that a
# single configured rate can populate two fields the backtest engine charges separately. Read the
# total, not the halves.
#
# What the curve measures, once it exists, is turnover. Cost enters the return series through
# `|delta w|` at each rebalance, so a strategy's sensitivity to the assumed rate is set by how
# much of the book it moves and how often, not by how good its predictions are. Two strategies
# with the same gross Sharpe can have breakevens that differ by an order of magnitude, and the
# whole reason to plot a curve rather than report one number is that the difference is invisible
# at any single rate. The quantity to read off is where the curve crosses zero and how far that
# sits from the quoted band above - a strategy that survives to 40 basis points on pairs that
# trade at 3 has room; one that dies at 4 is reporting an edge that is really a spread.

# %% tags=["results"]
cost_grid = get_cost_grid_bps(CASE_STUDY_ID)
if MAX_COST_POINTS:
    cost_grid = cost_grid[:MAX_COST_POINTS]
if not cost_grid:
    raise RuntimeError("the cost grid is empty")


def _catalog_row(result: BacktestResult) -> pl.DataFrame:
    prediction_hash = result.registry_record()["prediction_hash"]
    row = catalog.filter(pl.col("prediction_hash") == prediction_hash)
    if row.height != 1:
        raise RuntimeError(f"prediction {prediction_hash} resolved to {row.height} catalog rows")
    return row


def _strategy_arguments(result: BacktestResult) -> dict[str, Any]:
    strategy = result.spec()["strategy"]
    return {
        "signal": deepcopy(strategy["signal"]),
        "allocation": deepcopy(strategy.get("allocation")),
        "risk": deepcopy(strategy.get("risk")),
        "execution_mode": strategy.get("rebalance", {}).get("mode"),
    }


def _non_cost_projection(spec: dict[str, Any]) -> dict[str, Any]:
    projected = deepcopy(spec)
    projected.pop("chapter", None)
    projected.pop("_runtime_backtest_config", None)
    config = projected.get("backtest_config", {})
    config.pop("commission", None)
    config.pop("slippage", None)
    metadata = config.get("metadata")
    if isinstance(metadata, dict):
        metadata.pop("chapter", None)
        # An absolute filesystem path, and `case_studies/utils/registry/specs.py` already
        # excludes it from the identity hash for that reason. Comparing it here made the
        # check fail on where the notebook was run from rather than on what it produced:
        # a sibling written in one worktree never matches a parent registered in another,
        # and the message says a strategy field moved when none did.
        metadata.pop("preset_path", None)
    # The parent was serialized by whatever engine version registered it and the sibling by the
    # installed one, so a field the engine has since added to `BacktestConfig` is present on one
    # side and absent on the other while both describe the same strategy. `ml4t-backtest` 0.1.3 to
    # 0.1.6 added `account.lock_notional_update_mode` and `position_sizing.share_rounding`, which
    # failed every fx_pairs and crypto_perps_funding parent registered before 2026-09-12 with a
    # message saying a strategy field moved. Round-tripping both sides through the installed schema
    # states the comparison in one vocabulary, so the check answers what this notebook built rather
    # than which engine wrote the row it is compared against - and it covers the next added field
    # without naming it. `ensure_backtest_spec` deliberately does NOT round-trip, because there the
    # result is hashed and dropped unknown keys would move an identity; here it is compared and
    # discarded. Metadata is merged back over the serialized view for the same reason it is there:
    # the dataclass pins a schema and drops keys it does not know.
    if EngineBacktestConfig is not None and config:
        original_metadata = dict(metadata) if isinstance(metadata, dict) else {}
        rebuilt = EngineBacktestConfig.from_dict(config).to_dict()
        rebuilt.pop("commission", None)
        rebuilt.pop("slippage", None)
        rebuilt_metadata = dict(rebuilt.get("metadata") or {})
        rebuilt_metadata.update(original_metadata)
        rebuilt["metadata"] = rebuilt_metadata
        projected["backtest_config"] = rebuilt
    return projected


cost_jobs = []
for label, selected in selected_by_label.items():
    arguments = _strategy_arguments(selected)
    for total_bps in cost_grid:
        costs = {
            "commission_bps": total_bps / 2.0,
            "slippage_bps": total_bps / 2.0,
        }
        plan = plan_backtests(
            study,
            predictions=_catalog_row(selected),
            signal=arguments["signal"],
            allocation=arguments["allocation"],
            risk=arguments["risk"],
            costs=costs,
            chapter="ch18",
            execution_mode=arguments["execution_mode"],
        )
        if len(plan.members) != 1:
            raise RuntimeError("a cost plan must contain exactly one backtest")
        cost_jobs.append(
            {
                "label": label,
                "selected": selected,
                "arguments": arguments,
                "total_bps": total_bps,
                "costs": costs,
                "backtest_hash": plan.expected_hashes[0],
            }
        )

planned_hashes = [job["backtest_hash"] for job in cost_jobs]
if len(planned_hashes) != len(set(planned_hashes)):
    raise RuntimeError("two planned cost requests collapse to the same identity")

cost_population = None
if not include_preview:
    costs_name = research_name(CASE_STUDY_ID, "cost-sensitivity-backtests", scope=POPULATION_NAME)
    cost_population = OfficialPopulation.create(
        study,
        name=costs_name,
        member_kind="backtest",
        members=planned_hashes,
        supersedes=population_supersedes(
            study, name=costs_name, declared=SUPERSEDES_COST_BACKTESTS
        ),
    )
    print(f"Frozen expected cost population: {cost_population.hash}")

# %% [markdown]
# ## Execute the frozen cost grid, and validate its membership without making it selectable
#
# The population is validated in the cell that fills it: the expected set was written down before
# the first member ran, and `require_complete` is what turns that declaration into a published
# result. Publishing it does not make it selectable - a cost sensitivity is a curve through a
# parameter the strategy does not choose, and later selection reads the allocation population.
#
# "Registered but not selectable" is a distinction worth being concrete about, because both parts
# are deliberate. These rows are written to the registry, complete and current, exactly like every
# other backtest: the curve is a published result that a reader can look up and re-derive. What
# makes them unselectable is that the downstream stages read named populations rather than
# querying the registry for whatever is complete, so a cost sibling is never a member of a set
# anything ranks. The separation lives in which population a stage reads, not in a flag on the
# row - which is why a stage that queried the registry directly would silently acquire eleven
# copies of one strategy, each at a different assumed rate, and would rank them.

# %% tags=["results"]
# A sweep that recomputes everything and a sweep that recomputes nothing print the same summary
# unless the two are counted apart. `run_backtests` serves an identity that is already registered
# and complete instead of running it again, which is what makes a re-run affordable and what makes
# a bare member count say nothing about whether this run did any work.
#
# The runner already knows which it did and says so per member in `execution.diagnostics`, as
# `status` "reused" or "completed". Comparing against the registered hashes instead would be
# wrong in both directions: a registered-but-partial backtest is in that set, gets recomputed and
# would report as reused, and a preview re-run reads a table that excludes preview rows by default
# and would report every reused member as computed.
run_status: list[str] = []
cost_results: list[BacktestResult] = []
cost_rows = []
for job in cost_jobs:
    selected = job["selected"]
    arguments = job["arguments"]
    execution = run_backtests(
        study,
        predictions=_catalog_row(selected),
        signal=arguments["signal"],
        allocation=arguments["allocation"],
        risk=arguments["risk"],
        costs=job["costs"],
        chapter="ch18",
        execution_mode=arguments["execution_mode"],
    )
    if len(execution.results) != 1:
        raise RuntimeError("a cost request must produce exactly one backtest")
    result = execution.results[0]
    if result.hash != job["backtest_hash"]:
        raise RuntimeError("a completed cost identity differs from the frozen plan")
    if _non_cost_projection(result.spec()) != _non_cost_projection(selected.spec()):
        raise RuntimeError("a cost sibling changed a non-cost strategy field")
    if result.registry_record()["stage"] != "cost_sensitivity" or not result.complete:
        raise RuntimeError("a cost result is incomplete or misclassified")
    cost_results.append(result)
    run_status.extend(entry["status"] for entry in execution.diagnostics)
    cost_rows.append(
        {
            "label": job["label"],
            "total_cost_bps": job["total_bps"],
            "backtest_hash": result.hash,
            "prediction_hash": result.registry_record()["prediction_hash"],
        }
    )

served = run_status.count("reused")
print(
    f"Cost siblings: {reuse_disclosure(len(cost_results) - served, served)}, "
    f"{len(cost_results)} in the population"
)

if not include_preview:
    if cost_population is None:
        raise RuntimeError("the canonical cost population was not frozen before execution")
    cost_population.require_complete()
    print(f"Official cost-sensitivity population: {cost_population.hash}")
else:
    print("Preview cost curves remain outside official populations and candidate sets.")

pl.DataFrame(cost_rows).sort("label", "total_cost_bps")

# %% [markdown]
# ## Key takeaways
#
# - Each label has one validation-selected parent strategy, chosen across all three upstream
#   stages so the curve prices the strategy the chapter reports rather than a sibling of it.
# - Cost siblings preserve every non-cost identity field.
# - Cost sensitivity is frozen for completeness but excluded from later selection. A strategy that
#   could compete on its cost assumption would win by assuming costs away.
# - Basis points are the FX regime because spreads are quoted as a fraction of the rate. The
#   declared components are a taxonomy; the backtest charges one aggregate rate per traded leg.
# - The curve is a turnover measurement. Read where it crosses zero and compare that to the
#   quoted band, not the Sharpe at any single rate.
#
# The breakeven this produces is still a validation-period number, and turnover is not stable
# across regimes: a strategy that trades more in volatile periods pays more exactly when spreads
# are widest, and a curve computed at a constant rate cannot show that. The curve bounds the
# question rather than settling it.

```

Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT

Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.