رفتن به محتوا
همه اسناد کتابخانه

آزمون پوشش‌های ریسک برای استراتژی‌های آتی CME

نوت‌بوک یادگیری ماشین برای معامله‌گری

خلاصه

این دفترچه قواعد حد ضرر، حد ضرر متحرک و خروج زمانیِ پیکربندی‌شده را برای قوی‌ترین سیگنال یا استراتژی تخصیصِ رتبه‌بندی‌شده بر اساس اعتبارسنجی، برای هر افق بازده آتی اعمال می‌کند. هر قاعده جداگانه روی استراتژی پایه خودش ارزیابی می‌شود و پارامترها از پیش تثبیت شده‌اند، نه اینکه بر اساس مسیر قیمت اعتبارسنجی تنظیم شوند. همه قواعد تعریف‌شده باید کامل اجرا شوند تا جمعیت نامزدهای موردنظر برای انتخاب بعدی حفظ شود.

بحث توضیح می‌دهد که حدهای توقف با خروج نامتقارن از موقعیت‌ها، توزیع بازده استراتژی را تغییر می‌دهند. در معامله انتقالیِ بازگشت‌به‌میانگین، حد توقف ممکن است زمانی خروج را تحمیل کند که موقعیت بر اساس سیگنال جذاب‌ترین است؛ بنابراین نمی‌توان فرض کرد کاهش ریسک نتیجه را بهتر می‌کند. پوشش‌های منفرد جداگانه ارزیابی می‌شوند؛ اثرشان را نمی‌توان با هم جمع کرد، زیرا چند قاعده ممکن است با یک حرکت قیمت از موقعیت خارج شوند. نامزدهای پوشش ریسک همچنان می‌توانند در انتخاب نهایی اعتبارسنجی قرار گیرند، اما پوشش‌ها ممکن است خروج و ورود دوباره بیفزایند؛ بنابراین هزینه‌های معامله برای سنجش مفیدبودن بهبود ظاهری شارپ اهمیت دارند. نتایج شواهد اعتبارسنجی‌اند، نه اثبات منفعت خارج از نمونه.

ایده‌های کلیدی

  • پوشش ریسک زمان خروج استراتژی موجود از موقعیت‌ها را تغییر می‌دهد، نه اینکه تعیین کند چه چیزی نگه داشته شود.
  • پارامترهای حد ضرر، حد ضرر متحرک و خروج زمانی پیش از اعتبارسنجی تثبیت می‌شوند تا با مسیر قیمت آن تنظیم نشوند.
  • حدهای توقف ممکن است با سیگنال‌های معامله انتقالیِ بازگشت‌به‌میانگین تعارض داشته باشند، چون پس از زیان‌هایی که ممکن است پیش از بهبود رخ دهند، موقعیت را می‌بندند.
  • هر پوشش جداگانه در برابر استراتژی پایه خودش آزموده می‌شود؛ بنابراین بهبودهای مجزا ثابت نمی‌کنند که قواعد در ترکیب خوب عمل می‌کنند.
  • خروج‌ها و ورودهای دوباره اضافی می‌توانند باعث شوند هزینه‌های معامله، بهبود ظاهری اعتبارسنجی را معکوس کنند.

برچسب‌ها

متن کامل
# CME Futures: Risk Overlays


# CME Futures: Risk Overlays

For each return horizon, this notebook selects the highest validation Sharpe from the immutable
union of signal and allocation results, then applies every position-level risk rule declared in
the case-study configuration. Stop-loss, trailing-stop, and time-exit parameters are fixed before
the validation backtest. They are not calibrated from the same validation price path they assess.

Risk rules execute inside the existing futures engine after product-keyed target decisions cross
the typed boundary. Every declared rule must finish, and the resulting per-label candidate sets
remain eligible for final validation selection.

## What a risk overlay is, and why it is a separate stage

The stages before this one decided *what to hold*: a signal ranked the products, an allocation
rule decided how much of each. A risk overlay decides *when to stop holding it* - it sits on
top of an existing set of positions and closes them on a condition the signal never
considered.

The three rules here are the standard family. A **stop-loss** exits when a position has lost
more than a set amount from entry. A **trailing stop** exits when it has given back a set
amount from its best level, so it protects an unrealized gain rather than only the entry
price. A **time exit** closes after a fixed holding period whatever the position is doing, on
the reasoning that a signal with a horizon has nothing to say beyond it.

### The asymmetry these introduce, which is the point and the danger

A signal is symmetric about its own prediction: it is as willing to be wrong in one direction
as the other. A stop is not. It truncates the loss side of the distribution and leaves the
gain side alone, and that is why it appeals.

What it also does is convert an unrealized loss into a realized one at the worst available
moment, and give up any recovery that would have followed. For a mean-reverting signal - which
describes carry, the signal this case study trades - that is a direct conflict: the position
is exited precisely when the thing the signal is betting on has become most attractive. A stop
on a mean-reverting strategy is not a free reduction in risk. It is a change to the strategy,
and it can easily be a change for the worse.

That is the whole reason this is measured rather than assumed. Risk management is the part of
a strategy where intuition is least reliable and where "obviously prudent" is applied without
testing more often than anywhere else in the pipeline.

### Why the parameters are fixed before the backtest, and not after

The stop distances and holding periods come from `config/setup.yaml` and are fixed before the
validation backtest runs. They are deliberately **not** calibrated on the price path they are
then assessed against.

The reason is that this stage is unusually easy to cheat at without noticing. Choosing a stop
level by trying several and keeping the one with the best validation Sharpe would find the
level that best avoided the particular drawdowns that particular history happened to contain,
and would report the result as a risk improvement. Nothing about it would generalize, and
nothing in the output frame would show what happened - the returns would simply look better.
Fixing the parameters in configuration is what makes the comparison between overlay and no
overlay a real one.

### Why every declared rule must finish

A rule that failed and was skipped would leave a candidate set that silently means "the rules
that happened to work", and the selection downstream would then choose from it as though it
were the declared set. Failing the notebook is the only outcome that keeps the population
equal to the configuration.

### These candidates stay eligible

Unlike the cost sweep, risk-overlay results are part of the final selection pool.
`19_strategy_analysis` selects over the union of signal, allocation and risk-overlay
backtests, so an overlay that genuinely improves validation Sharpe can be what the case study
ships - and one that does not is visible as such next to the configuration it was applied to.

```python
"""Run the declared CME futures risk-overlay population."""

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    create_label_candidate_sets,
    open_study,
    pre_overlay_results,
    product_universe_table,
    rank_by_validation_sharpe,
    run_official_backtest_requests,
    strategy_request_frame,
)
from case_studies.research.population import supersedes_for_run
from case_studies.utils.sweep_config import get_position_risk_controls, get_top_n_predictions
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []

# The risk population is immutable under its name, so a run whose members have moved has to say
# which generation it retires. Anything upstream that changes a backtest identity moves them - a
# corrected label, a changed accounting field, a re-run after a registry reset - and
# `OfficialPopulation.create` refuses to write a different member list under a name that already
# exists. Declared as a literal so that running the committed notebook as it stands recomputes
# the population on record. Empty for a first snapshot.
RISK_POPULATION = "cme_futures-risk-validation-v1"
SUPERSEDES_RISK_POPULATION: str = ""

# How many parents per label the overlay grid sits on. `None` reads
# `backtest.sweep.top_n_predictions.risk_overlay`, which every case study declares as 1, and one
# is narrow on purpose: an overlay is a second search over the same validation folds, so the
# question the book asks is whether a control improves the configuration the funnel already
# chose. Until 2026-09-20 that 1 was a literal `[0]` below rather than a number read from the
# declaration, which left this case study unable to answer at any other width while four others
# could. A run at a wider width changes the member list of every name this notebook publishes,
# so it needs its own `RISK_POPULATION` and `SUPERSEDES_CANDIDATE_SETS` the same way a narrowed
# run does.
TOP_N_COMBOS = None

# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "cme_futures-risk-fwd_ret_5d-v1": "live",
    "cme_futures-risk-fwd_ret_21d-v1": "live",
}
```

## Fixed per-label inputs and risk rules

No candidate cap or runtime-dependent skip is allowed. The configured list is the population.

**What the overlay is applied to.** An overlay needs an existing strategy to sit on, and there
is one per label: the highest validation Sharpe from the immutable union of the signal and
allocation stages. Taking the best of the two stages rather than the signal alone matters,
because an overlay applied to a weaker parent would be measuring the overlay against a
strategy the case study would not have shipped anyway.

**Why per label rather than one overall.** Each return horizon is a different prediction
problem and its best configuration is chosen within its own horizon. Picking one parent across
all labels would let the strongest horizon's configuration stand in for horizons it was never
fitted for, and the overlay comparison would then be confounded by which label the parent came
from. Every configured rule runs against every label's own parent, so the comparison within a
label is like for like.

**Why the configured list is the population, with no cap.** A runtime cap would make the set
depend on how long the run took, which means a re-run could select from a different set and
nothing would record that it had. The rules are declared in configuration precisely so the
population is a property of the configuration rather than of the execution.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
    if PREVIEW_LABELS:
        raise ValueError("canonical execution cannot declare preview reductions")
    labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
    if WORKSPACE is None or not PREVIEW_LABELS:
        raise ValueError("preview execution requires WORKSPACE and PREVIEW_LABELS")
    unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
    if unknown:
        raise ValueError(f"preview labels this case study does not declare: {unknown}")
    labels = tuple(PREVIEW_LABELS)
else:
    raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
if TOP_N_COMBOS is None:
    TOP_N_COMBOS = get_top_n_predictions("cme_futures", "risk_overlay")
if TOP_N_COMBOS < 1:
    raise ValueError("the risk overlay needs at least one parent per label")
universe = product_universe_table()
universe
```

```python
risk_controls = get_position_risk_controls("cme_futures")
if not risk_controls:
    raise ValueError("the configured position-risk population is empty")

request_rows = []
for label in labels:
    ranked = rank_by_validation_sharpe(
        study,
        pre_overlay_results(
            study,
            label=label,
            execution_tier=EXECUTION_TIER,
            supersedes_by_set=SUPERSEDES_CANDIDATE_SETS,
        ),
    )
    for selected in ranked[:TOP_N_COMBOS]:
        strategy = selected.spec()["strategy"]
        prediction_hash = selected.registry_record()["prediction_hash"]
        for control in risk_controls:
            rule = {key: value for key, value in control.items() if key != "name"}
            request_rows.append(
                {
                    "request_name": f"{selected.hash}-risk-{control['name']}",
                    "prediction_hash": prediction_hash,
                    "label": label,
                    "signal": strategy["signal"],
                    "allocation": strategy.get("allocation"),
                    "risk": {"position_rules": [rule]},
                    "costs": None,
                    "chapter": "ch19",
                }
            )
requests = strategy_request_frame(request_rows)
requests.select("request_name", "prediction_hash", "label", "risk")
```

## Execute and freeze risk candidates

Each request carries the fitted prediction checkpoint, product decisions, fold-transition policy,
contract and roll inputs, and one risk rule. Missing members fail before the candidate set exists.

One request is one rule applied to one parent, so a rule's effect is read against its own
parent rather than against the field. Two rules that both improve Sharpe are not therefore
combinable: they may exit on the same moves, and their joint effect is not the sum of their
separate ones. Nothing here estimates that, and a reader stacking rules on the strength of
this table would be assuming an additivity it does not measure.

The results are frozen as a named population, and `SUPERSEDES_RISK_POPULATION` in the
parameter cell is how a re-run names the generation it retires. A retired snapshot stays in
the registry rather than being deleted, so a Sharpe quoted from an earlier generation remains
traceable to the population it was computed over.

```python
execution = run_official_backtest_requests(
    study,
    requests,
    population_name=RISK_POPULATION if EXECUTION_TIER == "canonical" else None,
    supersedes=supersedes_for_run(
        study,
        population_name=RISK_POPULATION,
        declared=SUPERSEDES_RISK_POPULATION or None,
        execution_tier=EXECUTION_TIER,
    ),
)
candidate_sets = (
    create_label_candidate_sets(
        study, execution, stage="risk", supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
    )
    if EXECUTION_TIER == "canonical"
    else {}
)
```

`source` says whether each member was computed by this run or served from the registry because
an identical identity was already recorded. A re-run of a registered sweep is entirely `reused`
and completes in seconds; without the column that is indistinguishable from having computed
every row.

```python
execution.catalog_rows.sort("label", "request_name")
```

Final selection in `19_strategy_analysis` uses the union of signal, allocation, and risk-overlay
results. Cost-sensitivity rows are excluded.

One consequence to carry into `16_costs`: an overlay only ever adds trades. Every stop that
fires is an exit that the signal did not ask for, and often a re-entry afterwards. So an
overlay that improves Sharpe here can still be the worse strategy once friction is priced, and
the two notebooks have to be read together rather than in sequence. This is also why the
selected configuration is priced with its overlay in place rather than bare.

با ذکر منبع و مطابق مجوز اثر، به‌طور کامل نمایش داده می‌شود. مجوز: MIT

این خلاصه را عامل پژوهشی Stratmill بر پایه متن اصلی نوشته است؛ نسخه‌ای از اثر منبع نیست.