Chuyển đến nội dung
Tất cả tài liệu trong thư viện

Kiểm tra lớp phủ rủi ro cho chiến lược hợp đồng tương lai CME

Notebook Machine Learning for Trading

Tóm tắt

Sổ tay áp dụng các quy tắc cắt lỗ, cắt lỗ động và thoát lệnh theo thời gian đã cấu hình cho tín hiệu hoặc chiến lược phân bổ có thứ hạng xác thực cao nhất ở mỗi chân trời lợi suất hợp đồng tương lai. Mỗi quy tắc được đánh giá riêng trên chiến lược nền tương ứng, với tham số cố định từ trước thay vì tinh chỉnh theo đường giá xác thực. Mọi quy tắc đã khai báo đều phải chạy xong để giữ nguyên quần thể ứng viên dự định dùng cho lựa chọn sau này.

Phần thảo luận giải thích rằng lệnh dừng thay đổi phân phối lợi suất của chiến lược bằng cách thoát vị thế bất đối xứng. Với chiến lược carry hồi quy về trung bình, lệnh dừng có thể buộc thoát khi vị thế hấp dẫn nhất theo tín hiệu, nên không thể mặc định rằng giảm rủi ro sẽ cải thiện kết quả. Các lớp phủ riêng lẻ được đánh giá tách biệt; không thể cộng tác động của chúng vì nhiều quy tắc có thể cùng thoát theo một biến động giá. Ứng viên có lớp phủ rủi ro vẫn đủ điều kiện được chọn cuối cùng trên dữ liệu xác thực, nhưng lớp phủ có thể tạo thêm lượt thoát và tái nhập, khiến chi phí giao dịch trở nên quan trọng khi đánh giá liệu mức cải thiện Sharpe biểu kiến có hữu ích hay không. Kết quả là bằng chứng xác thực, không phải bằng chứng về lợi ích ngoài mẫu.

Ý chính

  • Lớp phủ rủi ro thay đổi thời điểm chiến lược hiện có thoát vị thế, chứ không quyết định nắm giữ tài sản nào.
  • Tham số cắt lỗ, cắt lỗ động và thoát theo thời gian được cố định trước khi xác thực để tránh tinh chỉnh theo đường giá.
  • Lệnh dừng có thể xung đột với tín hiệu carry hồi quy về trung bình khi thoát sau khoản lỗ có thể được phục hồi.
  • Mỗi lớp phủ được thử riêng so với chiến lược nền tương ứng, nên lợi ích riêng lẻ không chứng minh các quy tắc kết hợp tốt.
  • Các lượt thoát và tái nhập bổ sung có thể khiến chi phí giao dịch đảo ngược mức cải thiện xác thực biểu kiến.

Thẻ

Toàn văn
# CME Futures: Risk Overlays


# CME Futures: Risk Overlays

For each return horizon, this notebook selects the highest validation Sharpe from the immutable
union of signal and allocation results, then applies every position-level risk rule declared in
the case-study configuration. Stop-loss, trailing-stop, and time-exit parameters are fixed before
the validation backtest. They are not calibrated from the same validation price path they assess.

Risk rules execute inside the existing futures engine after product-keyed target decisions cross
the typed boundary. Every declared rule must finish, and the resulting per-label candidate sets
remain eligible for final validation selection.

## What a risk overlay is, and why it is a separate stage

The stages before this one decided *what to hold*: a signal ranked the products, an allocation
rule decided how much of each. A risk overlay decides *when to stop holding it* - it sits on
top of an existing set of positions and closes them on a condition the signal never
considered.

The three rules here are the standard family. A **stop-loss** exits when a position has lost
more than a set amount from entry. A **trailing stop** exits when it has given back a set
amount from its best level, so it protects an unrealized gain rather than only the entry
price. A **time exit** closes after a fixed holding period whatever the position is doing, on
the reasoning that a signal with a horizon has nothing to say beyond it.

### The asymmetry these introduce, which is the point and the danger

A signal is symmetric about its own prediction: it is as willing to be wrong in one direction
as the other. A stop is not. It truncates the loss side of the distribution and leaves the
gain side alone, and that is why it appeals.

What it also does is convert an unrealized loss into a realized one at the worst available
moment, and give up any recovery that would have followed. For a mean-reverting signal - which
describes carry, the signal this case study trades - that is a direct conflict: the position
is exited precisely when the thing the signal is betting on has become most attractive. A stop
on a mean-reverting strategy is not a free reduction in risk. It is a change to the strategy,
and it can easily be a change for the worse.

That is the whole reason this is measured rather than assumed. Risk management is the part of
a strategy where intuition is least reliable and where "obviously prudent" is applied without
testing more often than anywhere else in the pipeline.

### Why the parameters are fixed before the backtest, and not after

The stop distances and holding periods come from `config/setup.yaml` and are fixed before the
validation backtest runs. They are deliberately **not** calibrated on the price path they are
then assessed against.

The reason is that this stage is unusually easy to cheat at without noticing. Choosing a stop
level by trying several and keeping the one with the best validation Sharpe would find the
level that best avoided the particular drawdowns that particular history happened to contain,
and would report the result as a risk improvement. Nothing about it would generalize, and
nothing in the output frame would show what happened - the returns would simply look better.
Fixing the parameters in configuration is what makes the comparison between overlay and no
overlay a real one.

### Why every declared rule must finish

A rule that failed and was skipped would leave a candidate set that silently means "the rules
that happened to work", and the selection downstream would then choose from it as though it
were the declared set. Failing the notebook is the only outcome that keeps the population
equal to the configuration.

### These candidates stay eligible

Unlike the cost sweep, risk-overlay results are part of the final selection pool.
`19_strategy_analysis` selects over the union of signal, allocation and risk-overlay
backtests, so an overlay that genuinely improves validation Sharpe can be what the case study
ships - and one that does not is visible as such next to the configuration it was applied to.

```python
"""Run the declared CME futures risk-overlay population."""

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    create_label_candidate_sets,
    open_study,
    pre_overlay_results,
    product_universe_table,
    rank_by_validation_sharpe,
    run_official_backtest_requests,
    strategy_request_frame,
)
from case_studies.research.population import supersedes_for_run
from case_studies.utils.sweep_config import get_position_risk_controls, get_top_n_predictions
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []

# The risk population is immutable under its name, so a run whose members have moved has to say
# which generation it retires. Anything upstream that changes a backtest identity moves them - a
# corrected label, a changed accounting field, a re-run after a registry reset - and
# `OfficialPopulation.create` refuses to write a different member list under a name that already
# exists. Declared as a literal so that running the committed notebook as it stands recomputes
# the population on record. Empty for a first snapshot.
RISK_POPULATION = "cme_futures-risk-validation-v1"
SUPERSEDES_RISK_POPULATION: str = ""

# How many parents per label the overlay grid sits on. `None` reads
# `backtest.sweep.top_n_predictions.risk_overlay`, which every case study declares as 1, and one
# is narrow on purpose: an overlay is a second search over the same validation folds, so the
# question the book asks is whether a control improves the configuration the funnel already
# chose. Until 2026-09-20 that 1 was a literal `[0]` below rather than a number read from the
# declaration, which left this case study unable to answer at any other width while four others
# could. A run at a wider width changes the member list of every name this notebook publishes,
# so it needs its own `RISK_POPULATION` and `SUPERSEDES_CANDIDATE_SETS` the same way a narrowed
# run does.
TOP_N_COMBOS = None

# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "cme_futures-risk-fwd_ret_5d-v1": "live",
    "cme_futures-risk-fwd_ret_21d-v1": "live",
}
```

## Fixed per-label inputs and risk rules

No candidate cap or runtime-dependent skip is allowed. The configured list is the population.

**What the overlay is applied to.** An overlay needs an existing strategy to sit on, and there
is one per label: the highest validation Sharpe from the immutable union of the signal and
allocation stages. Taking the best of the two stages rather than the signal alone matters,
because an overlay applied to a weaker parent would be measuring the overlay against a
strategy the case study would not have shipped anyway.

**Why per label rather than one overall.** Each return horizon is a different prediction
problem and its best configuration is chosen within its own horizon. Picking one parent across
all labels would let the strongest horizon's configuration stand in for horizons it was never
fitted for, and the overlay comparison would then be confounded by which label the parent came
from. Every configured rule runs against every label's own parent, so the comparison within a
label is like for like.

**Why the configured list is the population, with no cap.** A runtime cap would make the set
depend on how long the run took, which means a re-run could select from a different set and
nothing would record that it had. The rules are declared in configuration precisely so the
population is a property of the configuration rather than of the execution.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
    if PREVIEW_LABELS:
        raise ValueError("canonical execution cannot declare preview reductions")
    labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
    if WORKSPACE is None or not PREVIEW_LABELS:
        raise ValueError("preview execution requires WORKSPACE and PREVIEW_LABELS")
    unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
    if unknown:
        raise ValueError(f"preview labels this case study does not declare: {unknown}")
    labels = tuple(PREVIEW_LABELS)
else:
    raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
if TOP_N_COMBOS is None:
    TOP_N_COMBOS = get_top_n_predictions("cme_futures", "risk_overlay")
if TOP_N_COMBOS < 1:
    raise ValueError("the risk overlay needs at least one parent per label")
universe = product_universe_table()
universe
```

```python
risk_controls = get_position_risk_controls("cme_futures")
if not risk_controls:
    raise ValueError("the configured position-risk population is empty")

request_rows = []
for label in labels:
    ranked = rank_by_validation_sharpe(
        study,
        pre_overlay_results(
            study,
            label=label,
            execution_tier=EXECUTION_TIER,
            supersedes_by_set=SUPERSEDES_CANDIDATE_SETS,
        ),
    )
    for selected in ranked[:TOP_N_COMBOS]:
        strategy = selected.spec()["strategy"]
        prediction_hash = selected.registry_record()["prediction_hash"]
        for control in risk_controls:
            rule = {key: value for key, value in control.items() if key != "name"}
            request_rows.append(
                {
                    "request_name": f"{selected.hash}-risk-{control['name']}",
                    "prediction_hash": prediction_hash,
                    "label": label,
                    "signal": strategy["signal"],
                    "allocation": strategy.get("allocation"),
                    "risk": {"position_rules": [rule]},
                    "costs": None,
                    "chapter": "ch19",
                }
            )
requests = strategy_request_frame(request_rows)
requests.select("request_name", "prediction_hash", "label", "risk")
```

## Execute and freeze risk candidates

Each request carries the fitted prediction checkpoint, product decisions, fold-transition policy,
contract and roll inputs, and one risk rule. Missing members fail before the candidate set exists.

One request is one rule applied to one parent, so a rule's effect is read against its own
parent rather than against the field. Two rules that both improve Sharpe are not therefore
combinable: they may exit on the same moves, and their joint effect is not the sum of their
separate ones. Nothing here estimates that, and a reader stacking rules on the strength of
this table would be assuming an additivity it does not measure.

The results are frozen as a named population, and `SUPERSEDES_RISK_POPULATION` in the
parameter cell is how a re-run names the generation it retires. A retired snapshot stays in
the registry rather than being deleted, so a Sharpe quoted from an earlier generation remains
traceable to the population it was computed over.

```python
execution = run_official_backtest_requests(
    study,
    requests,
    population_name=RISK_POPULATION if EXECUTION_TIER == "canonical" else None,
    supersedes=supersedes_for_run(
        study,
        population_name=RISK_POPULATION,
        declared=SUPERSEDES_RISK_POPULATION or None,
        execution_tier=EXECUTION_TIER,
    ),
)
candidate_sets = (
    create_label_candidate_sets(
        study, execution, stage="risk", supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
    )
    if EXECUTION_TIER == "canonical"
    else {}
)
```

`source` says whether each member was computed by this run or served from the registry because
an identical identity was already recorded. A re-run of a registered sweep is entirely `reused`
and completes in seconds; without the column that is indistinguishable from having computed
every row.

```python
execution.catalog_rows.sort("label", "request_name")
```

Final selection in `19_strategy_analysis` uses the union of signal, allocation, and risk-overlay
results. Cost-sensitivity rows are excluded.

One consequence to carry into `16_costs`: an overlay only ever adds trades. Every stop that
fires is an exit that the signal did not ask for, and often a re-entry afterwards. So an
overlay that improves Sharpe here can still be the worse strategy once friction is priced, and
the two notebooks have to be read together rather than in sequence. This is also why the
selected configuration is priced with its overlay in place rather than bare.

Hiển thị toàn văn kèm ghi nguồn theo giấy phép của tài liệu gốc. Giấy phép: MIT

Bản tóm tắt này do tác nhân nghiên cứu của Stratmill biên soạn từ tài liệu gốc; đây không phải bản sao của tài liệu.