Prueba de capas de riesgo en estrategias de futuros de CME
Resumen
Este cuaderno aplica reglas configuradas de stop-loss, stop dinámico y salida por tiempo a la estrategia de señales o de asignación mejor clasificada en validación para cada horizonte de rendimiento de futuros. Cada regla se evalúa con su propia estrategia de referencia y sus parámetros se fijan de antemano, en vez de ajustarse a la trayectoria de precios de validación. Todas las reglas declaradas deben ejecutarse para preservar la población de candidatos prevista para la selección posterior.
La discusión explica que los stops cambian la distribución de rendimientos de una estrategia al cerrar posiciones de forma asimétrica. En una estrategia de carry con reversión a la media, un stop puede forzar la salida cuando la posición resulta más atractiva según la señal; por tanto, no se puede suponer que reducir el riesgo mejore los resultados. Las capas individuales se evalúan por separado; sus efectos no se pueden sumar, porque varias reglas pueden provocar una salida ante los mismos movimientos de precios. Los candidatos con capas de riesgo siguen siendo aptos para la selección final en validación, pero estas capas pueden añadir salidas y reentradas, por lo que los costes de transacción son importantes para juzgar si una aparente mejora del Sharpe resulta útil. Los resultados son evidencia de validación, no una prueba de beneficio fuera de muestra.
Ideas clave
- Una capa de riesgo cambia cuándo una estrategia existente cierra posiciones, no qué activos elige mantener.
- Los parámetros de stop-loss, stop dinámico y salida por tiempo se fijan antes de la validación para evitar ajustarlos a su trayectoria de precios.
- Los stops pueden entrar en conflicto con las señales de carry con reversión a la media al cerrar posiciones tras pérdidas que podrían preceder a una recuperación.
- Cada capa se prueba por separado frente a su propia estrategia de referencia, por lo que las ganancias individuales no demuestran que las reglas se combinen bien.
- Las salidas y reentradas adicionales pueden hacer que los costes de transacción reviertan una aparente mejora en validación.
Etiquetas
Texto completo
# CME Futures: Risk Overlays
# CME Futures: Risk Overlays
For each return horizon, this notebook selects the highest validation Sharpe from the immutable
union of signal and allocation results, then applies every position-level risk rule declared in
the case-study configuration. Stop-loss, trailing-stop, and time-exit parameters are fixed before
the validation backtest. They are not calibrated from the same validation price path they assess.
Risk rules execute inside the existing futures engine after product-keyed target decisions cross
the typed boundary. Every declared rule must finish, and the resulting per-label candidate sets
remain eligible for final validation selection.
## What a risk overlay is, and why it is a separate stage
The stages before this one decided *what to hold*: a signal ranked the products, an allocation
rule decided how much of each. A risk overlay decides *when to stop holding it* - it sits on
top of an existing set of positions and closes them on a condition the signal never
considered.
The three rules here are the standard family. A **stop-loss** exits when a position has lost
more than a set amount from entry. A **trailing stop** exits when it has given back a set
amount from its best level, so it protects an unrealized gain rather than only the entry
price. A **time exit** closes after a fixed holding period whatever the position is doing, on
the reasoning that a signal with a horizon has nothing to say beyond it.
### The asymmetry these introduce, which is the point and the danger
A signal is symmetric about its own prediction: it is as willing to be wrong in one direction
as the other. A stop is not. It truncates the loss side of the distribution and leaves the
gain side alone, and that is why it appeals.
What it also does is convert an unrealized loss into a realized one at the worst available
moment, and give up any recovery that would have followed. For a mean-reverting signal - which
describes carry, the signal this case study trades - that is a direct conflict: the position
is exited precisely when the thing the signal is betting on has become most attractive. A stop
on a mean-reverting strategy is not a free reduction in risk. It is a change to the strategy,
and it can easily be a change for the worse.
That is the whole reason this is measured rather than assumed. Risk management is the part of
a strategy where intuition is least reliable and where "obviously prudent" is applied without
testing more often than anywhere else in the pipeline.
### Why the parameters are fixed before the backtest, and not after
The stop distances and holding periods come from `config/setup.yaml` and are fixed before the
validation backtest runs. They are deliberately **not** calibrated on the price path they are
then assessed against.
The reason is that this stage is unusually easy to cheat at without noticing. Choosing a stop
level by trying several and keeping the one with the best validation Sharpe would find the
level that best avoided the particular drawdowns that particular history happened to contain,
and would report the result as a risk improvement. Nothing about it would generalize, and
nothing in the output frame would show what happened - the returns would simply look better.
Fixing the parameters in configuration is what makes the comparison between overlay and no
overlay a real one.
### Why every declared rule must finish
A rule that failed and was skipped would leave a candidate set that silently means "the rules
that happened to work", and the selection downstream would then choose from it as though it
were the declared set. Failing the notebook is the only outcome that keeps the population
equal to the configuration.
### These candidates stay eligible
Unlike the cost sweep, risk-overlay results are part of the final selection pool.
`19_strategy_analysis` selects over the union of signal, allocation and risk-overlay
backtests, so an overlay that genuinely improves validation Sharpe can be what the case study
ships - and one that does not is visible as such next to the configuration it was applied to.
```python
"""Run the declared CME futures risk-overlay population."""
from case_studies.cme_futures.research_workflow import (
ALL_LABELS,
create_label_candidate_sets,
open_study,
pre_overlay_results,
product_universe_table,
rank_by_validation_sharpe,
run_official_backtest_requests,
strategy_request_frame,
)
from case_studies.research.population import supersedes_for_run
from case_studies.utils.sweep_config import get_position_risk_controls, get_top_n_predictions
```
```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []
# The risk population is immutable under its name, so a run whose members have moved has to say
# which generation it retires. Anything upstream that changes a backtest identity moves them - a
# corrected label, a changed accounting field, a re-run after a registry reset - and
# `OfficialPopulation.create` refuses to write a different member list under a name that already
# exists. Declared as a literal so that running the committed notebook as it stands recomputes
# the population on record. Empty for a first snapshot.
RISK_POPULATION = "cme_futures-risk-validation-v1"
SUPERSEDES_RISK_POPULATION: str = ""
# How many parents per label the overlay grid sits on. `None` reads
# `backtest.sweep.top_n_predictions.risk_overlay`, which every case study declares as 1, and one
# is narrow on purpose: an overlay is a second search over the same validation folds, so the
# question the book asks is whether a control improves the configuration the funnel already
# chose. Until 2026-09-20 that 1 was a literal `[0]` below rather than a number read from the
# declaration, which left this case study unable to answer at any other width while four others
# could. A run at a wider width changes the member list of every name this notebook publishes,
# so it needs its own `RISK_POPULATION` and `SUPERSEDES_CANDIDATE_SETS` the same way a narrowed
# run does.
TOP_N_COMBOS = None
# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
"cme_futures-risk-fwd_ret_5d-v1": "live",
"cme_futures-risk-fwd_ret_21d-v1": "live",
}
```
## Fixed per-label inputs and risk rules
No candidate cap or runtime-dependent skip is allowed. The configured list is the population.
**What the overlay is applied to.** An overlay needs an existing strategy to sit on, and there
is one per label: the highest validation Sharpe from the immutable union of the signal and
allocation stages. Taking the best of the two stages rather than the signal alone matters,
because an overlay applied to a weaker parent would be measuring the overlay against a
strategy the case study would not have shipped anyway.
**Why per label rather than one overall.** Each return horizon is a different prediction
problem and its best configuration is chosen within its own horizon. Picking one parent across
all labels would let the strongest horizon's configuration stand in for horizons it was never
fitted for, and the overlay comparison would then be confounded by which label the parent came
from. Every configured rule runs against every label's own parent, so the comparison within a
label is like for like.
**Why the configured list is the population, with no cap.** A runtime cap would make the set
depend on how long the run took, which means a re-run could select from a different set and
nothing would record that it had. The rules are declared in configuration precisely so the
population is a property of the configuration rather than of the execution.
```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
if PREVIEW_LABELS:
raise ValueError("canonical execution cannot declare preview reductions")
labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
if WORKSPACE is None or not PREVIEW_LABELS:
raise ValueError("preview execution requires WORKSPACE and PREVIEW_LABELS")
unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
if unknown:
raise ValueError(f"preview labels this case study does not declare: {unknown}")
labels = tuple(PREVIEW_LABELS)
else:
raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
if TOP_N_COMBOS is None:
TOP_N_COMBOS = get_top_n_predictions("cme_futures", "risk_overlay")
if TOP_N_COMBOS < 1:
raise ValueError("the risk overlay needs at least one parent per label")
universe = product_universe_table()
universe
```
```python
risk_controls = get_position_risk_controls("cme_futures")
if not risk_controls:
raise ValueError("the configured position-risk population is empty")
request_rows = []
for label in labels:
ranked = rank_by_validation_sharpe(
study,
pre_overlay_results(
study,
label=label,
execution_tier=EXECUTION_TIER,
supersedes_by_set=SUPERSEDES_CANDIDATE_SETS,
),
)
for selected in ranked[:TOP_N_COMBOS]:
strategy = selected.spec()["strategy"]
prediction_hash = selected.registry_record()["prediction_hash"]
for control in risk_controls:
rule = {key: value for key, value in control.items() if key != "name"}
request_rows.append(
{
"request_name": f"{selected.hash}-risk-{control['name']}",
"prediction_hash": prediction_hash,
"label": label,
"signal": strategy["signal"],
"allocation": strategy.get("allocation"),
"risk": {"position_rules": [rule]},
"costs": None,
"chapter": "ch19",
}
)
requests = strategy_request_frame(request_rows)
requests.select("request_name", "prediction_hash", "label", "risk")
```
## Execute and freeze risk candidates
Each request carries the fitted prediction checkpoint, product decisions, fold-transition policy,
contract and roll inputs, and one risk rule. Missing members fail before the candidate set exists.
One request is one rule applied to one parent, so a rule's effect is read against its own
parent rather than against the field. Two rules that both improve Sharpe are not therefore
combinable: they may exit on the same moves, and their joint effect is not the sum of their
separate ones. Nothing here estimates that, and a reader stacking rules on the strength of
this table would be assuming an additivity it does not measure.
The results are frozen as a named population, and `SUPERSEDES_RISK_POPULATION` in the
parameter cell is how a re-run names the generation it retires. A retired snapshot stays in
the registry rather than being deleted, so a Sharpe quoted from an earlier generation remains
traceable to the population it was computed over.
```python
execution = run_official_backtest_requests(
study,
requests,
population_name=RISK_POPULATION if EXECUTION_TIER == "canonical" else None,
supersedes=supersedes_for_run(
study,
population_name=RISK_POPULATION,
declared=SUPERSEDES_RISK_POPULATION or None,
execution_tier=EXECUTION_TIER,
),
)
candidate_sets = (
create_label_candidate_sets(
study, execution, stage="risk", supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
)
if EXECUTION_TIER == "canonical"
else {}
)
```
`source` says whether each member was computed by this run or served from the registry because
an identical identity was already recorded. A re-run of a registered sweep is entirely `reused`
and completes in seconds; without the column that is indistinguishable from having computed
every row.
```python
execution.catalog_rows.sort("label", "request_name")
```
Final selection in `19_strategy_analysis` uses the union of signal, allocation, and risk-overlay
results. Cost-sensitivity rows are excluded.
One consequence to carry into `16_costs`: an overlay only ever adds trades. Every stop that
fires is an exit that the signal did not ask for, and often a re-entry afterwards. So an
overlay that improves Sharpe here can still be the worse strategy once friction is priced, and
the two notebooks have to be read together rather than in sequence. This is also why the
selected configuration is priced with its overlay in place rather than bare.Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.