مواد پر جائیں
لائبریری کی تمام دستاویزات

عمل درآمد، اخراجات اور رن شناخت کے لیے بیک ٹیسٹ پری سیٹس

کوڈ Machine Learning for Trading

خلاصہ

یہ کوڈ بیان کرتا ہے کہ کیس اسٹڈی کی ترتیب سے بیک ٹیسٹ کی ترتیبات کیسے طے ہوتی ہیں اور رن کی شناخت ٹریڈ کیے گئے مجموعے اور پیش گوئی کی عمر میں تبدیلیوں کو کیسے شمار کرتی ہے۔ بنیادی سبق یہ ہے کہ پورٹ فولیو یا پیش گوئیوں کا مجموعہ بدلنے والی ہر چھانٹی حکمتِ عملی کی تفصیلات میں درج ہونی چاہیے؛ ورنہ مختلف رنز کا ہیش ایک جیسا ہو سکتا ہے اور وہ ایک دوسرے کے نتائج دوبارہ استعمال کر سکتے ہیں۔ کائنات کی شناخت ترتیب دی گئی علامتوں کے ڈائجسٹ سے ظاہر ہوتی ہے، جبکہ عمر کا اعلان حد، اور اس سے ہٹائی گئی قطاروں اور علامتوں کا ریکارڈ رکھتا ہے۔

پری سیٹ بنانے والا حصہ اعلان کردہ تاخیر کو اسی بار یا اگلی بار کے عمل درآمد سے جوڑتا ہے، ڈیٹا کی تعدد اخذ کرتا ہے، اور اکاؤنٹ، کمیشن اور سلپیج کی ترتیبات مقرر کرتا ہے۔ یہ فی حصص کمیشن اور اسپریڈ پر مبنی سلپیج بھی سپورٹ کرتا ہے، جس میں صحیح عدد کے حصص کے حجم کا تعین اور جب عمل درآمد پہلے ہی کوٹ کی طرف کی قیمتیں استعمال کرے تو اسپریڈ کا دوبارہ چارج روکنا شامل ہے۔ کوڈ ایک ناپی گئی مثال دیتا ہے جس میں شناخت کی نگرانی شامل کرنے سے پہلے محدود اور مکمل کائنات پر کارکردگی یکساں تھی، اور پیش گوئی کی مطابقت کی مثال جس میں عمر کی حد نے قطاریں ہٹا دیں۔ یہ مثالیں حفاظتی تدابیر کی وجہ واضح کرتی ہیں مگر مذکورہ ڈیٹا سیٹس تک مخصوص ہیں؛ کوڈ خود بنیادی ڈھانچہ ہے، منافع بخش ٹریڈنگ اسٹریٹیجی کا ثبوت نہیں۔

اہم خیالات

  • کم کیے گئے ٹریڈنگ مجموعے کو رن کی شناخت میں ظاہر کریں تاکہ ایک علامتی مجموعے کا نتیجہ دوسرے کے لیے دوبارہ استعمال نہ ہو۔
  • پیش گوئی کی عمر کے اعلان میں اصل ہٹائی گئی قطاریں درج ہونی چاہئیں، کیونکہ ری فٹ ایک ہی عمر کی حد کے تحت مختلف پیش گوئی مجموعے بنا سکتے ہیں۔
  • فِل ٹائمنگ کے ٹوکن واضح اسی بار یا اگلی بار کے عمل درآمد طریقوں سے جڑتے ہیں۔
  • فی حصص کمیشن اور اسپریڈ سلپیج ایک واحد فیصد شرح سے زیادہ تفصیلی لاگت کے مفروضے فراہم کرتے ہیں۔
  • جب کوٹ کی طرف عمل درآمد میں اسپریڈ پہلے ہی چارج ہو تو اسی اسپریڈ کو دوبارہ سلپیج میں شامل کرنا دہری گنتی ہوگی۔

ٹیگز

مکمل متن
# backtest_presets.py


```py
from __future__ import annotations

from copy import deepcopy
from pathlib import Path
from typing import Any

import polars as pl
import yaml

from case_studies.utils.backtest_loaders import (
    VECTORIZED_CASE_STUDIES,
    declared_rebalance_step,
    declares_rebalance_step,
)
from case_studies.utils.backtest_loaders import BacktestConfig as CaseStudyBacktestConfig
from utils.paths import get_case_study_dir

try:
    from ml4t.backtest import BacktestConfig as EngineBacktestConfig
except (ImportError, ModuleNotFoundError):  # pragma: no cover - import depends on env
    EngineBacktestConfig = None


# The token names the bar a signal is filled on, not the price within it. Every
# entry here that maps to "next_bar" fills on the bar after the decision bar; which
# price that is belongs to the dataset. `NEXT_SESSION_CLOSE` is `sp500_options`,
# whose AlgoSeek option chain carries one end-of-session quote per contract per day
# and no open, so the next bar's only price is its close.
_EXECUTION_MODE_BY_DELAY = {
    "NEXT_BAR_OPEN": "next_bar",
    "NEXT_SESSION_CLOSE": "next_bar",
    "MONDAY_OPEN": "next_bar",
    "1_BAR": "next_bar",
    "AT_FUNDING_TIMESTAMP": "same_bar",
}


def traded_universe_declaration(prices: pl.DataFrame) -> dict[str, Any]:
    """Describe the symbol set a run can hold, for the caller to put in its spec.

    ``MAX_SYMBOLS`` reduces the price panel and nothing else. On the engine path that
    bounds what the run can trade, because the engine cannot fill an unpriced name; on
    the vectorized path it bounds nothing, because ``gross_ret = weight * y_true`` is
    computed from the predictions and ``prices`` supplies only the rebalance calendar.
    Either way the reduction never reached ``backtest_hash``, so a reduced run and a
    full run over the same predictions hashed alike and the second was served the
    first's result. Measured on us_firm_characteristics/11_backtest, 2026-08-24: 8
    predictions x 4 schemes at 300 symbols and at 3,708 gave bit-identical Sharpe, CAGR
    and drawdown across all 32 backtests.

    The declaration carries a digest of the sorted symbol list rather than its length,
    because ``{A, B}`` and ``{A, C}`` are two symbols each and two different portfolios.
    It goes in ``strategy.signal``, which is hashed whole, so a caller that declares one
    gets an identity of its own and can neither be served nor serve a full-universe row.
    ``run_backtest`` reads it back through ``apply_traded_universe``: it checks the panel
    it was handed against this digest and narrows the predictions to it, so the run
    trades what its identity says on both paths.

    Call it only when the panel was deliberately reduced. A caller that declares nothing
    produces byte-identical specs to before this key existed, which is what leaves every
    registered backtest at the identity it was written under.
    """
    from case_studies.utils.registry.specs import canonical_json, compute_hash

    if "symbol" not in prices.columns:
        raise ValueError(
            "traded_universe_declaration needs a 'symbol' column on the price panel; "
            f"got columns={list(prices.columns)}"
        )
    symbols = sorted(prices.get_column("symbol").drop_nulls().unique().to_list())
    if not symbols:
        raise ValueError("traded_universe_declaration was handed an empty price panel")
    return {
        "n_symbols": len(symbols),
        "digest": compute_hash(canonical_json({"symbols": symbols})),
    }


def prediction_age_declaration(
    *, max_age_sessions: int, dropped: int, kept: int, symbols_dropped: int
) -> dict[str, Any]:
    """Describe a prediction-age bound that removed rows, for the caller to put in its spec.

    The sibling of :func:`traded_universe_declaration`, and the same defect it exists for.
    Aligning predictions onto a coarser bar with a backward as-of join reuses a symbol's last
    score for as long as the panel runs after its series ends, and bounding that age changes
    what the backtest is computed from while ``prediction_hash`` and the strategy spec stay
    identical - so a run with the bound and a run without it hash alike, and the second is
    served the first's result. Measured on `nasdaq100_microstructure/17_costs`: the bound
    removes 19,667 of 303,641 aligned rows at 30-minute cadence.

    Carries the counts and not only the bound, because two runs can declare the same bound
    and drop different rows: the prediction set behind them moves on every refit, and a run
    whose universe left earlier is a different portfolio at the same tolerance.

    Declare it only when the bound actually removed something. A caller that drops nothing
    passes ``None`` and produces byte-identical specs to before this key existed, which is
    what leaves every registered backtest at the identity it was written under.
    """
    if dropped <= 0:
        raise ValueError(
            "prediction_age_declaration is for a bound that removed rows; a bound that "
            "removed none must declare nothing, so the run keeps the identity it had."
        )
    return {
        "max_age_sessions": int(max_age_sessions),
        "dropped": int(dropped),
        "kept": int(kept),
        "symbols_dropped": int(symbols_dropped),
    }


def resolve_execution_mode(fill_timing: str):
    """Map fill_timing string to ExecutionMode enum.

    Raises ValueError for unknown tokens instead of silently degrading.
    """
    from ml4t.backtest import ExecutionMode

    token = fill_timing.upper().replace(" ", "_")
    mode_str = _EXECUTION_MODE_BY_DELAY.get(token)
    if mode_str is None:
        raise ValueError(
            f"Unknown execution delay '{fill_timing}'. "
            f"Known values: {sorted(_EXECUTION_MODE_BY_DELAY.keys())}"
        )
    return ExecutionMode.NEXT_BAR if mode_str == "next_bar" else ExecutionMode.SAME_BAR


def preset_path(case_study: str) -> Path:
    """Path to the source-controlled backtest preset.

    Always reads from the source repo (never ML4T_OUTPUT_DIR), since
    config/backtest/base.yaml is checked-in source, not runtime data.
    """
    from utils.paths import get_case_study_source_dir

    return get_case_study_source_dir(case_study) / "config" / "backtest" / "base.yaml"


def load_backtest_preset(case_study: str) -> dict[str, Any]:
    path = preset_path(case_study)
    with path.open() as f:
        data = yaml.safe_load(f) or {}
    if not isinstance(data, dict):
        raise TypeError(f"Backtest preset at {path} must be a mapping")
    return data


def _infer_data_frequency(cadence: str) -> str:
    token = cadence.lower()
    if "15" in token:
        return "15m"
    if "30" in token:
        return "30m"
    if "1_hour" in token or "hourly" in token or token == "1h":
        return "1h"
    if "8_hour" in token or "funding" in token:
        return "irregular"
    return "daily"


def _build_feed_spec(
    case_study: str,
    prices: pl.DataFrame,
    case_config: CaseStudyBacktestConfig,
) -> dict[str, Any]:
    columns = set(prices.columns)
    feed = {
        "timestamp_col": "timestamp",
        "entity_col": "symbol",
        "open_col": "open" if "open" in columns else None,
        "high_col": "high" if "high" in columns else None,
        "low_col": "low" if "low" in columns else None,
        "close_col": "close" if "close" in columns else None,
        "price_col": "price" if "price" in columns else ("close" if "close" in columns else None),
        "volume_col": "volume" if "volume" in columns else None,
        "bid_col": "bid" if "bid" in columns else None,
        "ask_col": "ask" if "ask" in columns else None,
        "mid_col": "mid" if "mid" in columns else None,
        "calendar": case_config.calendar,
        "data_frequency": _infer_data_frequency(case_config.cadence),
        "timezone": "UTC",
    }
    if case_study in {"sp500_options", "sp500_equity_option_analytics"}:
        feed["bar_type"] = "quote"
    return {k: v for k, v in feed.items() if v is not None}


def build_resolved_backtest_config(
    case_study: str,
    case_config: CaseStudyBacktestConfig,
    strategy_spec: dict[str, Any],
    *,
    prices: pl.DataFrame,
    initial_cash: float,
) -> EngineBacktestConfig:
    if EngineBacktestConfig is None:  # pragma: no cover - import depends on env
        raise ImportError("ml4t-backtest is required for backtest preset resolution")

    preset = deepcopy(load_backtest_preset(case_study))
    preset.setdefault("account", {})
    preset.setdefault("execution", {})
    preset.setdefault("commission", {})
    preset.setdefault("slippage", {})
    preset.setdefault("cash", {})
    preset.setdefault("calendar", {})
    preset.setdefault("position_sizing", {})

    fill_timing = strategy_spec.get("execution", {}).get(
        "fill_timing"
    ) or case_config.execution_delay.upper().replace(" ", "_")
    execution_mode = _EXECUTION_MODE_BY_DELAY.get(fill_timing, "next_bar")
    signal_spec = strategy_spec.get("signal", {})
    signal_direction = str(signal_spec.get("direction", "long_only")).strip().lower()
    allow_short = bool(signal_spec.get("long_short", case_config.long_short)) or (
        signal_direction == "short_only"
    )

    preset["account"]["allow_short_selling"] = allow_short
    preset["execution"]["execution_mode"] = execution_mode
    preset["cash"]["initial"] = float(initial_cash)

    costs = strategy_spec.get("costs", {})
    cost_model = costs.get("model", "percentage")

    if cost_model == "percentage":
        commission_bps = float(costs.get("commission_bps", case_config.commission_bps))
        slippage_bps = float(costs.get("slippage_bps", case_config.slippage_bps))
        preset["commission"]["model"] = "percentage"
        preset["commission"]["rate"] = commission_bps / 10_000.0
        preset["slippage"]["model"] = "percentage"
        preset["slippage"]["rate"] = slippage_bps / 10_000.0
    elif cost_model == "per_share_plus_spread":
        # IB-style realistic equity costs: per-share commission, integer-share
        # sizing, half-spread slippage in dollars per share. Per-asset spreads
        # can be supplied directly (asset_spreads dict in setup.yaml) or via a
        # parquet artifact (asset_spreads_source) measured from quote data.
        preset["commission"]["model"] = "per_share"
        preset["commission"]["per_share"] = float(costs["per_share"])
        preset["commission"]["minimum"] = float(costs.get("minimum", 0.35))

        # When execution_price is quote_side, the fill price already includes
        # the bid/ask half-spread relative to mid (FillEngine returns
        # ask for BUY / bid for SELL via broker.QUOTE_SIDE). Wiring the same
        # measured half-spread into the slippage layer on top of that would
        # charge the spread twice. Use a zero slippage layer in that case;
        # per-share commission still applies independently.
        execution_price = preset.get("execution", {}).get("execution_price")
        if execution_price == "quote_side":
            preset["slippage"]["model"] = "percentage"
            preset["slippage"]["rate"] = 0.0
        else:
            preset["slippage"]["model"] = "spread"
            preset["slippage"]["spread_convention"] = costs.get("spread_convention", "half_spread")

            spread_by_asset: dict[str, float] = {}
            asset_spreads_source = costs.get("asset_spreads_source")
            if asset_spreads_source:
                spreads_path = get_case_study_dir(case_study, create=False) / asset_spreads_source
                spread_col = costs.get("asset_spreads_column", "median_half_spread_usd")
                spreads_df = pl.read_parquet(spreads_path)
                spread_by_asset = dict(
                    zip(
                        spreads_df["symbol"].to_list(),
                        [float(x) for x in spreads_df[spread_col].to_list()],
                    )
                )
            elif "asset_spreads" in costs:
                spread_by_asset = {str(k): float(v) for k, v in costs["asset_spreads"].items()}
            if spread_by_asset:
                preset["slippage"]["spread_by_asset"] = spread_by_asset
            if "default_half_spread_usd" in costs:
                preset["slippage"]["spread"] = float(costs["default_half_spread_usd"])
    else:
        raise ValueError(
            f"Unknown costs.model {cost_model!r}. Supported: 'percentage', 'per_share_plus_spread'."
        )

    # Share-type comes from setup.yaml::execution.share_type via case_config —
    # never hardcoded per-cost-model branch. Falls back to the preset JSON's
    # value (if any) when case_config.share_type is the default placeholder.
    share_type = getattr(case_config, "share_type", None)
    if share_type:
        preset["position_sizing"]["share_type"] = share_type

    preset["calendar"]["calendar"] = case_config.calendar
    preset["calendar"].setdefault("data_frequency", _infer_data_frequency(case_config.cadence))
    derived_feed = _build_feed_spec(case_study, prices, case_config)
    explicit_feed = preset.get("feed", {})
    preset["feed"] = {
        **derived_feed,
        **{k: v for k, v in explicit_feed.items() if v is not None},
    }

    metadata = dict(preset.get("metadata", {}))
    metadata.update(
        {
            "case_study": case_study,
            "chapter": strategy_spec.get("chapter"),
            "cadence": strategy_spec.get("execution", {}).get("cadence", case_config.cadence),
            "fill_timing": fill_timing,
            "preset_path": str(preset_path(case_study)),
            "signal_direction": signal_direction,
        }
    )
    preset["metadata"] = metadata

    return EngineBacktestConfig.from_dict(preset, preset_name=preset_path(case_study).stem)


def _serialize_backtest_config(config: EngineBacktestConfig | dict[str, Any]) -> dict[str, Any]:
    if hasattr(config, "to_dict"):
        return dict(config.to_dict())
    return dict(config)


def runtime_backtest_config(spec: dict[str, Any]) -> EngineBacktestConfig:
    if EngineBacktestConfig is None:  # pragma: no cover - import depends on env
        raise ImportError("ml4t-backtest is required for backtest config resolution")
    runtime = EngineBacktestConfig.from_dict(spec["backtest_config"])
    spec["_runtime_backtest_config"] = runtime
    return runtime


# Calendars that require session enforcement (drop bars outside trading
# sessions, e.g. CME Saturdays). Mirror of the rule in backtest_runner._run_engine;
# applied at spec construction so plan-time hashes match registered hashes.
SESSION_ENFORCED_CALENDARS = frozenset({"CME", "us_futures"})


def apply_calendar_session_enforcement(config: EngineBacktestConfig, calendar: str | None) -> None:
    """Set ``enforce_sessions=True`` on ``config`` when ``calendar`` requires it.

    Without this, ``ensure_backtest_spec`` would produce a plan-time spec
    with ``enforce_sessions=False`` while ``_run_engine`` later mutates the
    same runtime to ``True``, breaking ``_runtime_backtest_config`` hash
    stability.
    """
    if calendar in SESSION_ENFORCED_CALENDARS:
        config.enforce_sessions = True


def serializable_backtest_spec(spec: dict[str, Any]) -> dict[str, Any]:
    clean = deepcopy(spec)
    clean.pop("_runtime_backtest_config", None)
    if "backtest_config" in clean:
        clean["backtest_config"] = _serialize_backtest_config(clean["backtest_config"])
    return clean


def is_backtest_spec(spec: dict[str, Any]) -> bool:
    """Return True if ``spec`` is in canonical form (has ``strategy`` + ``backtest_config``)."""
    return spec.get("version") == 2 and "strategy" in spec and "backtest_config" in spec


def ensure_backtest_spec(
    case_study: str,
    case_config: CaseStudyBacktestConfig,
    strategy_spec: dict[str, Any],
    *,
    prices: pl.DataFrame,
    prediction_hash: str,
    initial_cash: float,
    traded_universe: dict[str, Any] | None = None,
) -> dict[str, Any]:
    """Normalize ``strategy_spec`` to the canonical backtest spec form.

    Idempotent: if ``strategy_spec`` is already canonical, it is returned with
    a refreshed ``_runtime_backtest_config``. Otherwise, a flat strategy_spec
    (with ``signal`` / ``execution`` / ``costs`` / etc. blocks) is projected
    into the canonical envelope: ``strategy.{signal, rebalance, allocation, risk}``
    plus a resolved ``backtest_config`` block.

    Rebalance thresholds (``min_weight_change``, ``min_trade_value``) are
    always populated in ``strategy.rebalance`` — taken from ``execution.*``
    when present, otherwise from ``case_config`` (which sources them from
    ``setup.yaml::backtest.rebalance.default``).

    ``traded_universe`` is the same declaration ``build_backtest_spec`` takes, for the
    same reason and with the same rule: build it with ``traded_universe_declaration``
    when the price panel was deliberately reduced, and it lands in ``strategy.signal``,
    which is hashed whole, so the reduction reaches ``backtest_hash`` instead of hashing
    like the full run over the same predictions. It is
    written only when a caller declares one, on both the carried-forward and the projected
    path, so every existing call site produces byte-identical specs and no registered
    backtest re-keys. A declaration handed here overwrites one inherited from the spec
    being normalized, because the panel this run was given is the universe it trades;
    ``apply_traded_universe`` then checks that claim against the panel in the runner.
    """
    if is_backtest_spec(strategy_spec):
        spec = deepcopy(strategy_spec)
        if case_study == "sp500_options":
            spec.setdefault("strategy", {}).setdefault("signal", {}).setdefault(
                "schedule_contract", SP500_OPTIONS_SCHEDULE_CONTRACT
            )
        spec.setdefault(
            "chapter", spec.get("backtest_config", {}).get("metadata", {}).get("chapter")
        )
        # Ensure rebalance thresholds are populated; specs that omit them
        # would otherwise raise KeyError on `rebalance_spec["min_weight_change"]`.
        rb = spec.setdefault("strategy", {}).setdefault("rebalance", {})
        rb.setdefault("min_weight_change", float(getattr(case_config, "min_weight_change", 0.005)))
        rb.setdefault("min_trade_value", float(getattr(case_config, "min_trade_value", 100.0)))
        if traded_universe is not None:
            spec["strategy"].setdefault("signal", {})["traded_universe"] = deepcopy(traded_universe)
        if "backtest_config" in spec:
            # Always overwrite metadata.prediction_hash with the caller's
            # argument — the spec may have been cloned from another run
            # (e.g. Ch20 holdout reuses the validation rank-1 spec with a
            # fresh holdout pred_hash), and downstream split-resolution
            # depends on the metadata reflecting the prediction set actually
            # being backtested.
            metadata = spec["backtest_config"].setdefault("metadata", {})
            metadata["prediction_hash"] = prediction_hash
            runtime = EngineBacktestConfig.from_dict(spec["backtest_config"])
            # Only the session-enforced calendars actually mutate ``runtime``
            # here; for all other calendars the from_dict/to_dict round-trip
            # would be a silent re-serialize that could perturb hashes for
            # any CS whose dict shape differs from the dataclass defaults
            # (None→0 normalization, dropped unknown keys, etc.). Confine the
            # ``backtest_config`` overwrite to the CSes where it is needed.
            if case_config.calendar in SESSION_ENFORCED_CALENDARS:
                apply_calendar_session_enforcement(runtime, case_config.calendar)
                # Re-serialize so the canonical ``backtest_config`` matches
                # the runtime; preserve every caller-supplied metadata key
                # by merging the original metadata back over the dataclass's
                # serialized view (the dataclass typically pins a schema
                # and drops unknown keys).
                original_metadata = dict(metadata)
                rebuilt = runtime.to_dict()
                rebuilt_metadata = dict(rebuilt.get("metadata") or {})
                rebuilt_metadata.update(original_metadata)
                rebuilt["metadata"] = rebuilt_metadata
                spec["backtest_config"] = rebuilt
            spec["_runtime_backtest_config"] = runtime
        return spec

    execution = deepcopy(strategy_spec.get("execution", {}))
    rebalance = {
        "mode": execution.get("mode", "engine"),
        "cadence": execution.get("cadence", case_config.cadence),
        "min_weight_change": float(
            execution["min_weight_change"]
            if "min_weight_change" in execution
            else getattr(case_config, "min_weight_change", 0.005)
        ),
        "min_trade_value": float(
            execution["min_trade_value"]
            if "min_trade_value" in execution
            else getattr(case_config, "min_trade_value", 100.0)
        ),
    }
    if "step" in execution:
        rebalance["step"] = int(execution["step"])
    strategy = {
        "signal": deepcopy(strategy_spec.get("signal", {})),
        "rebalance": rebalance,
    }
    if case_study == "sp500_options":
        strategy["signal"].setdefault("schedule_contract", SP500_OPTIONS_SCHEDULE_CONTRACT)
    if traded_universe is not None:
        strategy["signal"]["traded_universe"] = deepcopy(traded_universe)
    if "allocation" in strategy_spec:
        strategy["allocation"] = deepcopy(strategy_spec["allocation"])
    if "risk" in strategy_spec:
        strategy["risk"] = deepcopy(strategy_spec["risk"])

    resolved_config = build_resolved_backtest_config(
        case_study,
        case_config,
        strategy_spec,
        prices=prices,
        initial_cash=initial_cash,
    )
    resolved_config.metadata["prediction_hash"] = prediction_hash
    apply_calendar_session_enforcement(resolved_config, case_config.calendar)
    resolved_config_dict = resolved_config.to_dict()

    return {
        "version": 2,
        "chapter": strategy_spec.get("chapter"),
        "preset_id": f"{case_study}:base",
        "strategy": strategy,
        "backtest_config": resolved_config_dict,
        "_runtime_backtest_config": resolved_config,
    }


_COST_PASSTHROUGH_KEYS = (
    "model",
    "per_share",
    "minimum",
    "max_pct",
    "asset_spreads_source",
    "asset_spreads_column",
    "asset_spreads",
    "default_half_spread_usd",
    "spread_convention",
)

SP500_OPTIONS_SCHEDULE_CONTRACT = "last_available_session_per_iso_week_v1"


def _costs_block_from_case_config(
    case_config: CaseStudyBacktestConfig,
) -> dict[str, Any]:
    """Build the strategy_spec.costs block from the case config.

    For the percentage model (default), emit the bps form. For the
    per_share_plus_spread model, forward the full costs schema from setup.yaml
    so build_resolved_backtest_config can dispatch.
    """
    raw_costs = getattr(case_config, "raw_costs", None) or {}
    cost_model = raw_costs.get("model", "percentage")
    if cost_model == "per_share_plus_spread":
        return {key: deepcopy(raw_costs[key]) for key in _COST_PASSTHROUGH_KEYS if key in raw_costs}
    return {
        "commission_bps": case_config.commission_bps,
        "slippage_bps": case_config.slippage_bps,
    }


def build_backtest_spec(
    case_study: str,
    case_config: CaseStudyBacktestConfig,
    *,
    prices: pl.DataFrame,
    prediction_hash: str,
    initial_cash: float,
    signal: dict[str, Any],
    allocation: dict[str, Any] | None = None,
    risk: dict[str, Any] | None = None,
    costs: dict[str, Any] | None = None,
    chapter: str | None = None,
    execution_mode: str | None = None,
    min_weight_change: float | None = None,
    min_trade_value: float | None = None,
    label: str | None = None,
    traded_universe: dict[str, Any] | None = None,
    prediction_age: dict[str, Any] | None = None,
) -> dict[str, Any]:
    # A case study that declares per-label cadences must be told which label it is building for.
    # Defaulting to the case-study cadence here would put the spec on the wrong grid and register
    # it as if it were right, which is the failure this parameter exists to prevent - so it is a
    # refusal, not a fallback. Case studies that declare no override are unaffected.
    if getattr(case_config, "cadence_by_label", None) and not label:
        raise ValueError(
            f"{case_study} declares decision.cadence_by_label "
            f"({sorted(case_config.cadence_by_label)}) so build_backtest_spec needs a non-empty "
            "label=; "
            "pass the label this spec is being built for."
        )
    # The same refusal for the step, and for a sharper reason: the step is emitted below only
    # when a label resolves one, while `run_backtest` stamps the declared step onto whatever
    # spec it is handed. So an unlabelled call here does not build a spec on the wrong grid -
    # it builds a spec that hashes to an identity no run ever registers, and a caller that
    # pre-hashes it to decide what to skip finds every registered row missing
    # . Measured on us_firm_characteristics: 2,276 of 2,276
    # registered baseline rows invisible to `11_backtest`'s own skip check.
    if declares_rebalance_step(case_study) and not label:
        raise ValueError(
            f"{case_study} declares labels.rebalance_step, which is part of the backtest "
            "identity, so build_backtest_spec needs a non-empty label=; pass the label this "
            "spec is being built for. Without it the spec hashes to an identity no run "
            "registers."
        )
    resolved_signal = deepcopy(signal)
    if case_study == "sp500_options":
        resolved_signal.setdefault("schedule_contract", SP500_OPTIONS_SCHEDULE_CONTRACT)
    # The universe a reduced run trades, from `traded_universe_declaration`. It sits in the
    # signal block because that block is hashed whole, so the reduction reaches
    # `backtest_hash` without a second place to keep in step with it. Emitted only when the
    # caller declares one - the rule `cadence` and `step` already follow above - so a full
    # run produces byte-identical specs to before this parameter existed and every
    # registered backtest keeps the identity it was written under.
    if traded_universe is not None:
        resolved_signal["traded_universe"] = deepcopy(traded_universe)
    # The prediction-age bound, from `prediction_age_declaration`, on the same terms and in
    # the same block for the same reason: it changes what the run is computed from without
    # touching `prediction_hash`, so without this a bounded run and an unbounded one over the
    # same predictions hash alike and the second is served the first's result.
    if prediction_age is not None:
        resolved_signal["prediction_age"] = deepcopy(prediction_age)

    strategy_spec: dict[str, Any] = {
        "signal": resolved_signal,
        "execution": {
            # `VECTORIZED_CASE_STUDIES` rather than a second copy of the same set. The
            # membership was written out here as well, and the two are read by different
            # things: this one decides the dispatch in `backtest_runner`, and the constant
            # decides `IS_VECTORIZED` in four risk-management notebooks. Adding a case study
            # to one and not the other gives a notebook the wrong branch with nothing raising.
            "mode": (
                execution_mode
                if execution_mode is not None
                else "vectorized"
                if case_study in VECTORIZED_CASE_STUDIES
                else "engine"
            ),
            "engine_preset": "realistic",
            # Per-label when the case study declares one. `cadence_for` returns the case-study
            # default for an unlabelled caller and for a label with no override, so a case study
            # that declares nothing produces byte-identical specs to before this parameter existed.
            "cadence": case_config.cadence_for(label),
            # The step composes with the cadence to decide which slots are traded, so it
            # belongs to the identity beside it. Emitted only
            # when the case study declares one, so a case study that declares nothing
            # produces byte-identical specs to before this key existed - the same rule
            # `cadence_for` follows above.
            **(
                {"step": _step}
                if (_step := declared_rebalance_step(case_study, label)) is not None
                else {}
            ),
            "fill_timing": case_config.execution_delay.upper().replace(" ", "_"),
            "min_weight_change": (
                min_weight_change
                if min_weight_change is not None
                else getattr(case_config, "min_weight_change", 0.005)
            ),
            "min_trade_value": (
                min_trade_value
                if min_trade_value is not None
                else getattr(case_config, "min_trade_value", 100.0)
            ),
        },
        "costs": (
            deepcopy(costs) if costs is not None else _costs_block_from_case_config(case_config)
        ),
    }
    if chapter is not None:
        strategy_spec["chapter"] = chapter
    if allocation is not None:
        strategy_spec["allocation"] = deepcopy(allocation)
    if risk is not None:
        strategy_spec["risk"] = deepcopy(risk)
    return ensure_backtest_spec(
        case_study,
        case_config,
        strategy_spec,
        prices=prices,
        prediction_hash=prediction_hash,
        initial_cash=initial_cash,
    )


def clone_backtest_spec(spec: dict[str, Any]) -> dict[str, Any]:
    cloned = serializable_backtest_spec(spec)
    if is_backtest_spec(cloned):
        cloned["_runtime_backtest_config"] = EngineBacktestConfig.from_dict(
            cloned["backtest_config"]
        )
    return cloned


def set_backtest_costs_bps(
    spec: dict[str, Any],
    *,
    commission_bps: float,
    slippage_bps: float,
) -> dict[str, Any]:
    if not is_backtest_spec(spec):
        updated = deepcopy(spec)
        updated["costs"] = {
            "commission_bps": commission_bps,
            "slippage_bps": slippage_bps,
        }
        return updated

    updated = deepcopy(spec)
    bt_cfg = updated["backtest_config"]
    bt_cfg["commission"] = {
        "model": "percentage",
        "rate": commission_bps / 10_000.0,
    }
    bt_cfg["slippage"] = {
        "model": "percentage",
        "rate": slippage_bps / 10_000.0,
    }
    updated["_runtime_backtest_config"] = EngineBacktestConfig.from_dict(bt_cfg)
    return updated


def set_backtest_costs_per_share(
    spec: dict[str, Any],
    *,
    per_share: float,
    default_half_spread_usd: float,
    asset_spreads: dict[str, float] | None = None,
    spread_convention: str = "half_spread",
    minimum: float = 0.0,
) -> dict[str, Any]:
    """Mutate spec to use per-share commission + spread slippage.

    Switches the engine commission/slippage models from `percentage` to
    `per_share` / `spread`. Safe to call on a spec that originally used the
    percentage model — the previous rate fields are replaced. Used by the
    cost-sensitivity sweep to walk the per-share+spread regime alongside
    the bps regime for case studies whose dataset supports it (those with
    prices and integer-share semantics).

    `minimum` is the per-order commission floor in dollars and defaults to
    `0.0` (no floor). Pass `minimum=0.35` to match the IBKR Pro per-order
    floor used by `build_resolved_backtest_config`.
    """
    if not is_backtest_spec(spec):
        updated = deepcopy(spec)
        updated["costs"] = {
            "model": "per_share_plus_spread",
            "per_share": float(per_share),
            "default_half_spread_usd": float(default_half_spread_usd),
            "asset_spreads": dict(asset_spreads or {}),
            "spread_convention": spread_convention,
            "minimum": float(minimum),
        }
        return updated

    updated = deepcopy(spec)
    bt_cfg = updated["backtest_config"]
    bt_cfg["commission"] = {
        "model": "per_share",
        "rate": 0.0,
        "per_share": float(per_share),
        "minimum": float(minimum),
        "per_trade": 0.0,
    }
    bt_cfg["slippage"] = {
        "model": "spread",
        "rate": 0.0,
        "spread": float(default_half_spread_usd),
        "spread_convention": spread_convention,
    }
    if asset_spreads:
        bt_cfg["slippage"]["spread_by_asset"] = {str(k): float(v) for k, v in asset_spreads.items()}
    bt_cfg.setdefault("position_sizing", {})["share_type"] = "integer"
    updated["_runtime_backtest_config"] = EngineBacktestConfig.from_dict(bt_cfg)
    return updated


def strategy_view(spec: dict[str, Any]) -> dict[str, Any]:
    return spec["strategy"] if is_backtest_spec(spec) else spec


def cost_view(spec: dict[str, Any]) -> dict[str, Any]:
    if not is_backtest_spec(spec):
        return spec.get("costs", {})
    cfg = spec.get("backtest_config", {})
    commission = cfg.get("commission", {})
    slippage = cfg.get("slippage", {})
    return {
        "commission_bps": round(float(commission.get("rate", 0.0)) * 10_000.0, 10),
        "slippage_bps": round(float(slippage.get("rate", 0.0)) * 10_000.0, 10),
    }

```

ماخذ کا حوالہ دیتے ہوئے مکمل متن دکھایا گیا ہے، ماخذ کے لائسنس کے تحت۔ لائسنس: MIT

یہ خلاصہ اصل ماخذ سے Stratmill کے تحقیقی ایجنٹ نے لکھا ہے؛ یہ ماخذ کی نقل نہیں۔