Comparación de métodos de tamaño de posición en estrategias con futuros perpetuos de criptomonedas
Resumen
Este cuaderno estudia seis métodos de asignación aplicados a configuraciones seleccionadas de estrategias con futuros perpetuos de criptomonedas. Las ordenaciones y las reglas de entrada permanecen fijas mientras cambia el capital asignado a cada posición. Los métodos usan puntuaciones de modelos, volatilidad de cada contrato, covarianza entre contratos o incertidumbre histórica. Las configuraciones avanzan desde la referencia según el Sharpe del backtest de validación; las configuraciones de modelo distintas se cuentan por separado de sus conjuntos de predicciones. Cada asignación se compara con su propia referencia emparejada, manteniendo constantes las predicciones, el checkpoint, la regla de entrada, los costes y la financiación.
El cuaderno insiste en que el Sharpe más alto de una etapa no debe compararse directamente con el líder de la etapa anterior, ya que las configuraciones subyacentes son distintas. También distingue los asignadores que requieren una sola estimación de volatilidad de aquellos que estiman matrices de covarianza completas, lo que puede añadir ruido de estimación. Se conservan los resultados obtenidos al operar en particiones distintas, pero se excluyen de las comparaciones pareadas. La evidencia se basa en particiones de validación y una longitud fija de ventana móvil; en esta etapa también se fijan los costes, y la propia asignación puede afectar la rotación. Por tanto, las conclusiones no establecen el rendimiento holdout ni cuál es la mejor ventana o los mejores supuestos de costes.
Ideas clave
- La asignación cambia los pesos de las posiciones y mantiene fija la clasificación y la regla de entrada seleccionadas para cada una.
- Las configuraciones supervivientes se seleccionan por el Sharpe del backtest de validación, contando por separado las configuraciones de modelo distintas y los conjuntos de predicciones.
- Una comparación válida por etapas empareja cada resultado de asignación con su referencia correspondiente bajo los mismos supuestos.
- Los asignadores basados en covarianza estiman más cantidades que los métodos de volatilidad inversa y pueden ser más sensibles al ruido de estimación.
- Los resultados obtenidos al operar en particiones distintas no están emparejados y no deben leerse como comparaciones directas con la referencia.
Etiquetas
Texto completo
# Crypto perpetuals: six ways to size a position the ranking already chose
# Crypto perpetuals: six ways to size a position the ranking already chose
[`13_backtest`](13_backtest.ipynb) gave every position the same weight. That was the point
there - equal weight adds no information, so a difference between two baselines is a difference
between two rankings. This notebook keeps the rankings and the entry rules exactly where they
were and changes one thing: how much capital each admitted position gets.
Six alternatives are declared in `config/setup.yaml`, and they read three different kinds of
input. One reads the model's own score, so a contract the model is more confident about gets
more capital. Four read the history of returns - each contract's own volatility, or the
covariance between contracts - so that the positions contribute comparable amounts of risk
rather than comparable amounts of money. One reads how uncertain the model's prediction has
been on that contract in the past.
**This stage is narrow on purpose.** It runs on the survivors of the baseline rather than on
everything, because a sizing method applied to a ranking that lost money equally-weighted is
not a question anyone needs answered, and because every configuration added to a search makes
the highest result in it easier to reach by luck. The funnel that decides which baselines advance
is described below and is not a choice this notebook makes.
**Learning objectives.** By the end of this notebook you will be able to:
- Take the top model configurations from a completed baseline stage, counting distinct
configurations rather than distinct prediction sets, and say why the two differ.
- Say what each declared allocator reads, and which of them need a history of returns before
they can weight anything.
- Run a sizing variant so that it differs from its own baseline in one field, and read the
paired difference rather than comparing two leaders.
- Recognise when an allocator has traded a different set of dates from the one it is being
compared against, and why that makes the comparison an unpaired one.
**Book reference**: Chapter 17 (Portfolio Construction).
**Prerequisites**: [`13_backtest`](13_backtest.ipynb) has registered a complete
`stage='signal'` baseline for every declared prediction set.
**What it writes**: one `stage='allocation'` backtest per surviving prediction set, entry rule
and allocator, and one candidate set per label holding the baseline and the allocation results
together. [`15_risk_management`](15_risk_management.ipynb) and [`16_costs`](16_costs.ipynb)
read those.
```python
"""Run the declared allocator grid on the surviving crypto perpetuals baselines."""
import plotly.graph_objects as go
import polars as pl
from case_studies.crypto_perps_funding.research_workflow import (
ALL_LABELS,
)
from case_studies.research import (
CandidateSet,
Result,
candidate_set_supersedes,
open_study,
population_supersedes,
run_backtests,
)
from case_studies.research.strategy import strategy_warmup_periods
from case_studies.utils.backtest_loaders import load_backtest_prices_for
from case_studies.utils.registry.queries import resolve_best_predictions
from case_studies.utils.sweep_config import (
get_allocator_lookback,
get_allocators,
get_entry_schemes_for,
get_top_n_predictions,
)
from utils.artifact_specs import load_setup_config
from utils.style import COLORS, show_plotly_with_alt
```
```python
LABELS: list[str] = []
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
POPULATION_SUFFIX = "v1"
# None means the width `setup.yaml` declares; an int overrides it. Declared here because
# papermill only binds a name the parameters cell already holds - a run that passes
# TOP_N_PREDICTIONS to a notebook without it sweeps the declared width and exits 0.
TOP_N_PREDICTIONS = None
# Left empty, and it stays empty. The registry was reset for the stage-04 holdout rebuild, so
# every name below is published at generation one and there is nothing to supersede. A
# declaration is only needed when a re-run changes an existing name's membership: the refusal
# prints the name and the hash, and it is resolved through the shared resolver rather than
# offered straight, because a reader's clean clone has no generation for it to replace.
SUPERSEDES: dict[str, str] = {}
# Per-allocation-population lineage, keyed by the population's own name. `SUPERSEDES` above is
# keyed by label and covers the one per-label set frozen at the end; this stage declares one
# population per (label, scheme, allocator), so a single label key cannot name them. An entry
# for a name whose members did not change is ignored, so the whole current generation can be
# passed at once rather than discovered one refusal at a time.
# Left empty, and it stays empty. The registry was reset for the stage-04 holdout rebuild, so
# every name below is published at generation one and there is nothing to supersede. A
# declaration is only needed when a re-run changes an existing name's membership: the refusal
# prints the name and the hash, and it is resolved through the shared resolver rather than
# offered straight, because a reader's clean clone has no generation for it to replace.
SUPERSEDES_ALLOCATION: dict[str, str] = {}
```
```python
study = open_study(
"crypto_perps_funding", execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None
)
setup = load_setup_config("crypto_perps_funding")
labels = list(LABELS) if LABELS else list(ALL_LABELS)
# Where this run's own results are written and read back from: the released case directory on a
# canonical run, the isolated preview directory otherwise. `study.root` is the released one in
# both tiers, so a preview that reads it is reading somebody else's registry.
STORAGE_ROOT = study.storage_root(study.execution_tier)
n_assets = int(setup["universe"]["n_assets"])
# A canonical run publishes the funnel's named pools; a preview run reads and writes only its
# own results and names none of them. Both are the same sweep - the tier decides what is
# declared, never what is computed.
#
# The condition is the tier alone. It used to carry `and not WORKSPACE`, which read a canonical
# run into an isolated registry as a preview and sent it down the branch below that rebuilds the
# signal members from preview backtests this workspace has never held. A workspace is where the
# results go, not what they are: a candidate set is canonical either way, `CandidateSet.create`
# refuses a preview member, and a workspace seeded from the canonical registry carries exactly
# the frozen pools the canonical branch opens.
CANONICAL_RUN = EXECUTION_TIER == "canonical"
```
## 1. Which baselines advance
The selection funnel is sequential: each stage runs on the survivors of the one before it, and
the survivors are chosen on **validation backtest Sharpe** - never on information coefficient,
which measures whether a model ranks contracts correctly and says nothing about what a strategy
trading that ranking earns after costs and funding.
`config/setup.yaml` sets the width of each stage under `backtest.sweep.top_n_predictions`. The
baseline ran on everything; allocation runs on the top ten. **Ten what** is the part worth being
precise about: ten distinct `(family, config_name)` pairs, not ten prediction sets. A boosted
model contributes one prediction set per checkpoint, so counting prediction sets would let a
single configuration occupy the whole shortlist with ten readings of itself and crowd out every
other model in the case study. `checkpoints_per_config` then says how many checkpoints each
advancing configuration brings with it, and it is one - the checkpoint that scored best at the
baseline, since the checkpoint is part of the configuration rather than a knob to be re-tuned
here.
`resolve_best_predictions` is the implementation this case study uses, and it is not the only
one: five of the nine case studies call it, and the other four restate the same rule
separately - `cme_futures` in `research_workflow.shortlist_signal_configurations`, `fx_pairs`
in a function defined inside its own allocation notebook, `sp500_options` and
`us_equities_panel` inline over a ranked frame. They agree on what the rule is - rank on
validation Sharpe, keep the best row per `(family, config_name)`, take the declared width -
and differ in the pool they rank and in what they do when the pool is short of that width.
Nothing compares the four, so anything reasoning about which configurations advance without
running the notebook has to pick one and is wrong about the others
(ml4t/agent-workspace#1204).
**It is asked about one population, not about the registry.** `13_backtest` froze the baselines
each label hands on as `crypto-signal-{label}`, and that set is the whole of what may advance.
Ranking without it reads every signal backtest the registry holds, which includes the sweeps a
corrected re-run retired: those rows are immutable and their predictions are still current, so
no filter on the prediction side can exclude them, and `MAX(sharpe)` over a configuration then
ranks it on a number no live result carries. The set's members are backtest identities, and
`backtest_hashes` is where they go.
```python
if CANONICAL_RUN:
signal_members = {
label: set(CandidateSet.one(study, name=f"crypto-signal-{label}").members)
for label in labels
}
else:
# A preview run has no frozen set to open, because a candidate set is canonical. Its
# equivalent is the signal backtests its own 13_backtest wrote into this workspace.
preview_signal = study.backtests.table(include_preview=True).filter(
(pl.col("stage") == "signal")
& (pl.col("split") == "validation")
& (pl.col("execution_tier") == "preview")
& pl.col("label").is_in(labels)
)
if preview_signal.is_empty():
raise RuntimeError(
"no preview signal backtests in this workspace; 13_backtest has to run in the "
"same workspace before this notebook"
)
signal_members = {
label: set(
preview_signal.filter(pl.col("label") == label).get_column("backtest_hash").to_list()
)
for label in labels
}
if TOP_N_PREDICTIONS is None:
TOP_N_PREDICTIONS = get_top_n_predictions("crypto_perps_funding", "allocation")
top_n = TOP_N_PREDICTIONS
# `vertical_relaxed`, because `checkpoint_value` is Null-typed for a label whose survivors
# are all final-checkpoint models and Int64 for one that advanced a boosted model on a
# numbered checkpoint. Both are the same column meaning the same thing; a strict concat
# refuses the pair, and it refuses it only when more than one label is in play.
survivors = pl.concat(
[
resolve_best_predictions(
"crypto_perps_funding",
label,
split="validation",
stage="signal",
top_n=top_n,
checkpoints_per_config=1,
# The registry this run is writing to, which is the preview one on a preview run.
# `study.root` is the released case directory in both tiers, so passing it made a
# preview ask the canonical registry about backtests it never wrote there.
case_dir=study.storage_root(study.execution_tier),
backtest_hashes=signal_members[label],
)
for label in labels
],
how="vertical_relaxed",
)
if survivors.is_empty():
raise RuntimeError("no baseline survivors: run 13_backtest before this notebook")
```
One row per advancing configuration. `sharpe` is the baseline it advanced on, and it is shown
so the shortlist can be read against the population it came from rather than in isolation.
```python
catalog = study.predictions.table(include_preview=not CANONICAL_RUN).filter(
pl.col("prediction_hash").is_in(survivors.get_column("prediction_hash").implode())
)
if catalog.height != survivors.height:
raise RuntimeError("a surviving prediction is absent from the prediction catalog")
if catalog.filter(~pl.col("complete")).height:
raise RuntimeError("a surviving prediction set is incomplete")
survivors.select("label", "family", "config_name", "checkpoint_value", "sharpe").sort(
"label", "sharpe", descending=[False, True]
)
```
## 2. What each allocator reads
All six take the same set of admitted positions from the entry rule and decide only the size of
each. They differ in what they consult to do it.
- **`score_weighted`** is the only one that reads the prediction. Weight is proportional to the
model's score, so the ordering the ranking produced becomes a spread of position sizes rather
than a set of equal ones.
- **`inverse_vol`** weights each contract by the reciprocal of its own recent return
volatility. A contract that moves twice as much gets half the capital, so each position
contributes a comparable amount of risk. It looks at each contract on its own and ignores how
they move together.
- **`risk_parity`** equalizes each position's contribution to the volatility of the whole
portfolio, which requires the covariance between contracts rather than just the diagonal. On
a set of perpetual futures that mostly rise and fall together, the difference from
`inverse_vol` is the part correlation accounts for.
- **`hrp`**, hierarchical risk parity, groups the contracts by how correlated they are, splits
capital between the groups, and then splits it again inside each. It never inverts the
covariance matrix, which is what makes it usable when the estimate is noisy.
- **`mvo_ledoit_wolf`** does invert one: mean-variance optimization, on a covariance pulled part
of the way towards a simpler structured estimate. That pull is the **shrinkage**, and it is
there because a nineteen-by-nineteen covariance estimated from a few hundred observations
contains enough noise that optimizing against it directly concentrates the book on whichever
pair happens to look least correlated in the sample.
- **`conformal_weighted`** reads how wrong the model has been. For each contract it takes a
quantile of the model's own past absolute errors on that contract and gives less capital
where that quantile is larger. It is inverse-uncertainty sizing with a conformal quantile
standing in for a volatility estimate, and nothing here consumes it as an interval: the
weights are `1/width` normalized within each side at each timestamp, so any factor common
to every contract cancels and only the spread across contracts reaches the portfolio.
The four that read return history share one window, so that a difference between them is the
method rather than the amount of history each was given.
```python
allocators = get_allocators("crypto_perps_funding")
lookback = get_allocator_lookback("crypto_perps_funding")
bar_hours = int(setup["features"]["bar_hours"])
print(
f"Rolling window for the moment-based allocators: {lookback} bars of "
f"{bar_hours}-hourly prices, about {lookback * bar_hours / 24:.0f} days of history before "
"the first decision each can weight."
)
pl.DataFrame(
[
{
"allocator": allocator["method"],
"reads": "prediction score"
if allocator["method"] == "score_weighted"
else "prediction error history"
if allocator["method"] == "conformal_weighted"
else "return history",
"warmup_bars": strategy_warmup_periods({"allocation": allocator}),
}
for allocator in allocators
]
)
```
### What an allocator that calibrates costs, and what it does not
`conformal_weighted` needs residuals before it can size anything, and a residual is only usable
once the return it measures has been realized. So its width at a decision comes from every
error the model has already made on that contract up to the label horizon before that decision
- earlier folds and the current fold's own elapsed history alike. That costs a **warm-up**: the
first few decisions of the validation span have no width and the allocator holds nothing
through them. On this case study's 8-hourly grid the warm-up is three decisions.
It used to cost a great deal more. Calibrating on whole earlier folds only meant the earliest
fold had no earlier fold, so the allocator sat out the entire first year - and that is worth
keeping in view even though it is fixed, because of how the resulting number looked. **A period
a strategy sits out still counts as a period it observed.** A day holding nothing books a
return of exactly zero, so a book flat for a year reports the same period count as one that
traded every day of it, and every summary built on that count agreed the two were comparable.
The strategy that sat out 2022 posted the highest Sharpe in the stage by not trading a losing
year.
Section 4 therefore measures which folds each result actually **traded**, from its registered
return series, and pairs only within a matching set. That test is not about one allocator: any
result that covered part of the span for any reason is measured on a different sample from one
that covered all of it, and differencing the two attributes the sample to the sizing.
## 3. Running the grid
The entry rules are the ones the baseline established as feasible on a nineteen-contract
universe, unchanged, because changing the sizing and the selection together would measure
neither. For every surviving prediction set, every feasible entry rule and every allocator,
`run_backtests` resolves a strategy that differs from its baseline in the `allocation` field
and in nothing else.
Prices are loaded once per label and warmup. The moment-based allocators need the rolling
window of prices in front of their first decision and the other two do not, so there are two
price frames per label rather than one - and they are the frames the boundary would have
loaded for itself, so passing them in changes nothing but the number of reads.
```python
warmups = sorted({strategy_warmup_periods({"allocation": item}) for item in allocators})
prices_by_key = {
(label, warmup): load_backtest_prices_for(
"crypto_perps_funding", label, split="validation", warmup_periods=warmup
)
for label in labels
for warmup in warmups
}
schemes_by_label = {
label: get_entry_schemes_for("crypto_perps_funding", label, n_assets=n_assets, long_short=True)
for label in labels
}
```
```python
allocations = []
for label in labels:
label_rows = catalog.filter(pl.col("label") == label)
for scheme in schemes_by_label[label]:
signal = {key: value for key, value in scheme.items() if key != "name"}
for allocator in allocators:
warmup = strategy_warmup_periods({"allocation": allocator})
allocation_population = (
f"crypto-allocation-{label}-{scheme['name']}-"
f"{allocator['method']}-{POPULATION_SUFFIX}"
)
execution = run_backtests(
study,
predictions=label_rows,
signal=signal,
allocation=allocator,
prices=prices_by_key[(label, warmup)],
chapter="ch17",
population_name=allocation_population if CANONICAL_RUN else None,
# Resolved, not offered. The declaration below is committed source and is wrong
# on a reader's clean clone, where the registry holds no generation to supersede
# and `create` refuses a first version that claims to replace one.
supersedes=population_supersedes(
study,
name=allocation_population,
declared=SUPERSEDES_ALLOCATION.get(allocation_population),
)
if CANONICAL_RUN
else None,
)
allocations.append((label, scheme["name"], allocator["method"], execution))
print(
f"{label} / {scheme['name']} / {allocator['method']}: "
f"{len(execution.results)} backtests registered\n"
f" this execution: {execution.disclosure()}"
)
```
## 4. What came out
Read back from the registry, one row per allocator. `traded_folds` lists the validation folds
each result actually held a position in, derived below from the registered return series rather
than assumed from the allocator's name. A fold's number identifies it and does not order it -
the walk-forward split numbers folds as it builds them - so the list is written oldest first.
```python
# The rows are named: the baselines the frozen signal set admits, plus the allocation
# backtests this run just registered. Reading `stage IN (signal, allocation)` off the registry
# instead would fold every retired generation of both stages back into the grid, and a
# superseded result is not a candidate.
allocation_hashes = {
result.hash for _, _, _, execution in allocations for result in execution.results
}
in_play = set().union(*signal_members.values()) | allocation_hashes
results = study.backtests.table(include_preview=not CANONICAL_RUN).filter(
pl.col("backtest_hash").is_in(list(in_play))
)
if results.filter(~pl.col("complete")).height:
raise RuntimeError("the backtest catalog contains incomplete members")
entry_rule = (
pl.col("signal_method")
+ pl.when(pl.col("spec_json").str.json_path_match("$.strategy.signal.top_k").is_not_null())
.then(pl.lit("_top") + pl.col("spec_json").str.json_path_match("$.strategy.signal.top_k"))
.otherwise(pl.lit(""))
).alias("entry_rule")
keyed = results.with_columns(
entry_rule,
pl.col("allocation_method").fill_null("equal_weight").alias("allocator"),
)
```
Each decision carries its fold in the prediction set, and the dates a result held a
position on are in the return series it registered - a day the strategy was flat contributes a
return of exactly zero. Intersecting the two says which folds each result traded, which is the
property that has to match before two results can be differenced. The windows are put in date
order here because the fold numbers are not in date order.
```python
def fold_windows(label: str) -> pl.DataFrame:
"""First and last decision date of each validation fold, for one label.
Every configuration for a label predicts the same keys, which 13_backtest established, so
one prediction set carries the fold calendar for all of them.
"""
reference = catalog.filter(pl.col("label") == label).sort("prediction_hash")
return (
Result.open(study, reference.item(0, "prediction_hash"), include_preview=not CANONICAL_RUN)
.load()
.group_by("fold")
.agg(
fold_start=pl.col("timestamp").min().dt.date(),
fold_end=pl.col("timestamp").max().dt.date(),
)
.sort("fold_start")
)
```
```python
def traded_folds(backtest_hash: str, windows: pl.DataFrame) -> tuple[int, ...]:
"""Which validation folds one registered result actually held a position in."""
returns = pl.read_parquet(
STORAGE_ROOT / "run_log" / "backtest" / backtest_hash / "daily_returns.parquet"
)
column = next(name for name in returns.columns if name != "timestamp")
active = returns.filter(pl.col(column) != 0).select(pl.col("timestamp").dt.date().alias("day"))
if active.is_empty():
return ()
return tuple(
int(row["fold"])
for row in windows.iter_rows(named=True)
if active.filter(pl.col("day").is_between(row["fold_start"], row["fold_end"])).height
)
```
```python
windows_by_label = {label: fold_windows(label) for label in labels}
keyed = keyed.with_columns(
pl.Series(
"traded_folds",
[
"+".join(
str(fold)
for fold in traded_folds(row["backtest_hash"], windows_by_label[row["label"]])
)
for row in keyed.iter_rows(named=True)
],
)
)
```
```python
allocation_grid = (
keyed.filter(pl.col("stage") == "allocation")
.group_by("allocator")
.agg(
backtests=pl.len(),
labels=pl.col("label").n_unique(),
median_sharpe=pl.col("sharpe").median(),
above_zero=(pl.col("sharpe") > 0).sum(),
traded_folds=pl.col("traded_folds").unique().sort().str.join(", "),
median_turnover=pl.col("avg_turnover").median(),
)
.sort("allocator")
)
allocation_grid
```
### The paired difference, one field at a time
A leader-to-leader comparison across stages is not evidence that sizing helped: the allocation
leader and the baseline leader can be different models on different checkpoints, and the gap
between them then contains the search as well as the sizing. The join below pairs each
allocation result with the baseline built on the **same prediction set and the same entry
rule**, so the only field that differs is the allocator, and reports the difference.
A pair is admitted only when both legs traded the same folds. That is what excludes the
conformal rows, and it would exclude anything else that sat out part of the span for any other
reason - the test is what the result did, not which allocator produced it.
```python
baseline = keyed.filter(pl.col("stage") == "signal").select(
"prediction_hash",
"entry_rule",
pl.col("sharpe").alias("baseline_sharpe"),
pl.col("traded_folds").alias("baseline_traded_folds"),
)
allocation = keyed.filter(pl.col("stage") == "allocation")
paired = (
allocation.join(baseline, on=["prediction_hash", "entry_rule"], how="inner")
.filter(pl.col("traded_folds") == pl.col("baseline_traded_folds"))
.with_columns((pl.col("sharpe") - pl.col("baseline_sharpe")).alias("sharpe_change"))
)
unpaired = (
allocation.join(baseline, on=["prediction_hash", "entry_rule"], how="inner")
.filter(pl.col("traded_folds") != pl.col("baseline_traded_folds"))
.group_by("allocator")
.agg(left_unpaired=pl.len(), traded_folds=pl.col("traded_folds").unique().sort().str.join(", "))
.sort("allocator")
)
print(f"{paired.height} paired comparisons")
print(unpaired)
paired.group_by("allocator").agg(
pairs=pl.len(),
median_change=pl.col("sharpe_change").median(),
improved=(pl.col("sharpe_change") > 0).sum(),
best_change=pl.col("sharpe_change").max(),
worst_change=pl.col("sharpe_change").min(),
).sort("allocator")
```
One distribution per allocator, over every pair. The zero line is equal weight: a point above
it is a configuration that sizing improved, and the fraction of each distribution above the
line is the more informative reading than any single point in it.
```python
order = sorted(set(paired.get_column("allocator")))
fig = go.Figure()
for allocator in order:
panel = paired.filter(pl.col("allocator") == allocator)
fig.add_trace(
go.Box(
y=panel.get_column("sharpe_change").to_list(),
name=allocator,
marker_color=COLORS["blue"],
boxpoints="all",
jitter=0.4,
pointpos=0,
marker={"size": 3, "opacity": 0.5},
showlegend=False,
)
)
fig.add_hline(y=0, line_width=1, line_dash="dash", line_color=COLORS["neutral"])
fig.update_layout(
title={
"text": "Change in validation Sharpe from the equal-weight baseline"
"<br><sup>One point per prediction set and entry rule; the pair differs only in the "
"allocator</sup>",
"x": 0.02,
"xanchor": "left",
},
xaxis_title="Allocator",
yaxis_title="Sharpe minus its own equal-weight baseline",
height=520,
width=1000,
)
show_plotly_with_alt(
fig,
"Box plots with every pair overlaid as a point, one box per allocator, of the change in "
"annualized validation Sharpe against the equal-weight baseline built on the same prediction "
"set and entry rule. A dashed horizontal line marks zero, meaning no change from equal "
"weight. Every allocator's distribution straddles that line, and the boxes overlap one "
"another, so no allocator separates from equal weight or from the others.",
)
```
## 5. The candidate set each label hands on
The next two stages choose from the baseline and the allocation results **together**, which is
what the funnel prescribes: the question at the risk stage is whether an overlay helps the
highest-Sharpe configuration found so far, and equal weight is still eligible to be that
configuration. One set per label holds both stages.
**A candidate set admits only results that traded every validation fold**, and that exclusion
is doing real work rather than tidying. Selection downstream is on validation Sharpe, and a
Sharpe earned over one fold is not a larger or smaller version of one earned over two - it is a
measurement of a different period. A strategy that sits out a fold the others traded is
therefore not a stronger candidate when that fold went badly; it is an incomparable one, and
admitting it lets the choice of configuration turn on which period each candidate happened to
be exposed to.
For `conformal_weighted` this is structural rather than incidental: its intervals need an
earlier fold to calibrate on, so it can never trade the earliest one, and on a two-fold split
that is half the period. The results stay registered and visible in the grid above. What they
do not do is compete for a selection that would be reading exposure as skill.
A set is identified by its members, so a re-run that admits the same results returns the set
that already exists. A re-run that admits different ones - because something upstream was
corrected, or because the admission rule changed - is a second generation, and it has to name
the generation it replaces in `SUPERSEDES`. That is not ceremony: `15_risk_management` and
`19_strategy_analysis` both resolve this set by name, so two live
generations of one name would leave them unable to say which comparison a result came from.
The error raised on a changed set names the predecessor hash to pass.
```python
admitted_by_label: dict[str, list[str]] = {}
for label in labels:
label_rows = keyed.filter(pl.col("label") == label)
# Every fold the label declares, in date order - not the most common value observed, which
# would define full exposure as whatever the majority of results happened to reach.
full = "+".join(str(fold) for fold in windows_by_label[label].get_column("fold").to_list())
admitted = label_rows.filter(pl.col("traded_folds") == full)
excluded = label_rows.height - admitted.height
set_name = f"crypto-signal-allocation-{label}"
if CANONICAL_RUN:
members = study.backtests.freeze(
results.filter(
pl.col("backtest_hash").is_in(admitted.get_column("backtest_hash").implode())
),
name=set_name,
# Keyed by label, and also by the full set name, which is what the refusal prints.
# Pasting back the name it names is the obvious thing to try, and it used to miss.
# Resolved for the same reason the allocation lineage above is.
supersedes=candidate_set_supersedes(
study, name=set_name, declared=SUPERSEDES.get(set_name) or SUPERSEDES.get(label)
),
)
admitted_by_label[label] = list(members.members)
print(
f"{members.name}: {len(members.members)} members traded folds {full}; "
f"{excluded} excluded for trading fewer"
)
else:
admitted_by_label[label] = admitted.get_column("backtest_hash").to_list()
print(
f"{set_name} (preview): {len(admitted_by_label[label])} members traded folds "
f"{full}; {excluded} excluded for trading fewer, not frozen"
)
```
## 6. What to notice
**A sizing rule can only redistribute what the ranking selected.** None of the six changes
which contracts are held; they change how much of each. So the ceiling on what this stage can
add is set by the entry rule, and a ranking that selects the wrong contracts cannot be sized
into a good strategy. That is the reason the funnel puts sizing after selection rather than
searching the two together.
**The paired difference is the only honest reading of a stage increment.** Every allocation row
above has a baseline row built on the same prediction set, the same checkpoint, the same entry
rule, the same costs and the same funding, and the difference between those two numbers is the
allocator. The difference between this stage's highest Sharpe and the previous stage's highest
Sharpe is not: those are two different configurations, and most of the gap between them is the
search that produced them.
**Two allocators that look similar are doing different amounts of estimation.** `inverse_vol`
needs one number per contract; `risk_parity`, `hrp` and `mvo_ledoit_wolf` need a whole
covariance matrix, estimated from the same window. The more parameters an allocator estimates
from a fixed history, the more of its weights are noise, and on nineteen contracts with a few
hundred observations that is not a small consideration. It is also why the shrinkage in
`mvo_ledoit_wolf` and the clustering in `hrp` exist at all - both are ways of asking the same
data for fewer numbers.
**An unpaired row is reported, not quietly dropped.** A result that traded fewer folds than its
baseline is a real result and is registered like the others; what it is not is comparable to a
baseline that was holding positions while it was flat. Excluding such a row from the paired
frame while leaving it in the grid table is the distinction, and the frame printed beside the
count names which allocator it covers and which folds those rows actually traded. With every
allocator now calibrated on every fold, that frame is expected to be empty - which is what a
working guard looks like, not a reason to remove it.
**The count of observations is not the count of exposure.** A flat day still registers a return
of zero, so period counts, and any check built on them, agree exactly between a result that
traded a fold and one that sat it out. This is the kind of difference that a comparison hides
rather than reports, and finding it needs the return series rather than the summary row.
**Known limitations.** The rolling window is one length for every moment-based allocator, so
nothing here says whether a different amount of history would suit one of them better - that
would be another search axis, and adding it would widen the very search the funnel narrows.
Costs are the flat declared schedule, which sizing interacts with directly, since an allocator
that spreads capital more evenly turns over more of the book at each rebalance;
[`16_costs`](16_costs.ipynb) varies that assumption at the end of the funnel. And every number
is measured on the validation folds.
**Next**: [`15_risk_management`](15_risk_management.ipynb) holds the surviving configuration
fixed and asks whether a position-level control improves it. [`16_costs`](16_costs.ipynb) then
prices the winner, which is the one stage in the funnel that selects nothing.
Se muestra íntegramente con atribución según la licencia de la fuente. Licencia: MIT
Este resumen lo redactó el agente de investigación de Stratmill a partir del original; no es una copia de la fuente.