Comparação de alocadores de carteira para estratégias com opções do S&P 500
Resumo
Este notebook compara regras de ponderação de carteira para a venda de straddles em ativos selecionados do S&P 500. Mantém os ativos escolhidos e varia a alocação, usando pesos iguais como referência. Os métodos incluem ponderação por pontuação prevista, volatilidade inversa, paridade de risco, paridade hierárquica de risco, otimização média-variância com encolhimento e largura do intervalo de previsão conformal. Essas abordagens usam informações diferentes e podem concentrar capital ou estimar risco de maneiras distintas.
O processo de seleção fixa um conjunto de candidatos, classifica as referências pelo índice de Sharpe na validação, desempata pela identificação do backtest e conta configurações distintas de modelos em vez de linhas de checkpoint quase duplicadas. Comparações pareadas ajudam a atribuir diferenças à alocação mantendo fixos os demais campos da estratégia. Os resultados são estimativas pontuais de validação sem intervalos de incerteza e herdam o ruído de seleção da etapa de referência. Outra limitação é que os métodos baseados em covariância usam retornos das ações subjacentes para dimensionar posições em opções, tratando assim o risco do straddle de forma indireta; a ponderação conformal também omite datas sem intervalos calibrados, em vez de substituir silenciosamente outra regra.
Ideias principais
- Fixe o conjunto de configurações candidatas antes da classificação para que a regra de seleção use um universo de resultados reproduzível.
- Conte configurações distintas de modelos ao formar uma lista finalista para evitar que variantes de checkpoint ocupem o espaço de outros modelos.
- Mantenha os ativos e as demais configurações da estratégia fixos ao variar os alocadores, para facilitar a interpretação das comparações.
- A ponderação por pontuação acompanha a força das previsões, enquanto os métodos baseados em covariância buscam pesos ajustados ao risco a partir dos retornos subjacentes.
- Estimativas pontuais de validação herdam o ruído de seleção, e a covariância das ações subjacentes representa o risco do straddle apenas de forma indireta.
Tags
Texto completo
# 13_portfolio_management.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # S&P 500 Options: Portfolio Construction
#
# `12_backtest` weighted the straddles it sold equally: every symbol held on a decision date got
# the same share of capital. That is a deliberate null - it uses the model only to decide *which*
# symbols to trade, never *how much* of each. This notebook keeps the same symbols and varies the
# weighting rule, so that any difference in the result is attributable to the allocator and to
# nothing else.
#
# The rules come in two kinds. One reads the prediction itself and puts more capital behind a
# stronger score. The others ignore the prediction and read the covariance of the underlying
# returns, sizing positions so that each contributes comparable risk rather than comparable
# capital. Both kinds are common in practice and they fail in different ways, which is the point
# of running them side by side.
#
# The results extend the immutable candidate set that `18_strategy_analysis` selects from.
#
# **Learning objectives**
#
# - Freeze a set of finished backtests into a named candidate set whose membership cannot change
# afterwards, and read the selection rule off that set rather than off the registry.
# - Advance a fixed number of distinct model configurations to the next stage, counting
# configurations rather than backtest rows so that one model cannot occupy the shortlist.
# - Vary a single strategy field across an entire shortlist and keep every other field equal, so
# the comparison is paired.
#
# **Book reference**: Chapter 17
#
# **Prerequisites**: the complete baseline population published by `12_backtest`.
# %%
"""Execute the declared S&P 500 options allocation population."""
import plotly.express as px
import polars as pl
from case_studies.research import (
CandidateSet,
OfficialPopulation,
Result,
candidate_set_supersedes,
supersedes_for_run,
)
from case_studies.sp500_options.research_workflow import (
ALL_LABELS,
open_study,
paired_sharpe_on_common_support,
preview_baseline_candidates,
run_official_backtest_requests,
strategy_request_frame,
)
from case_studies.utils.sweep_config import (
get_allocators,
get_checkpoints_per_config,
get_top_n_predictions,
top_n_cap,
)
from utils.style import COLORS, show_plotly_with_alt
CASE_STUDY = "sp500_options"
BASELINE_POPULATION = "sp500-options-baseline-validation-v1"
BASELINE_CANDIDATES = "sp500-options-baseline-candidates-v1"
STRATEGY_CANDIDATES = "sp500-options-strategy-candidates-v1"
ALLOCATION_POPULATION = "sp500-options-allocation-validation-v1"
# %% tags=["parameters"]
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
PREVIEW_LABELS: list[str] = []
PREVIEW_MAX_BASELINE_CONFIGS = 0
PREVIEW_ALLOCATORS: tuple[str, ...] = ("score_weighted",)
# The generation each named set retires. A set and a population are immutable under their
# name, so a re-run whose membership moved has to say which one it replaces; the refusal
# names the current hash, and empty is correct only for a name this registry has never held.
# Each of these is stale the moment the run it authorizes succeeds, because that run becomes
# the generation the next one has to name.
SUPERSEDES_BASELINE_CANDIDATES: str = ""
SUPERSEDES_ALLOCATION_POPULATION: str = ""
SUPERSEDES_STRATEGY_CANDIDATES: str = ""
# None means the width `setup.yaml` declares; an int overrides it. Declared here because
# papermill only binds a name the parameters cell already holds - a run that passes
# TOP_N_PREDICTIONS to a notebook without it sweeps the declared width and exits 0.
TOP_N_PREDICTIONS = None
# %% [markdown]
# ## Freeze what is being selected from
#
# A candidate set is the list of results a selection is allowed to consider, written down and
# hashed before the selection happens. A selection rule only has a definite answer once the set it
# ranges over is fixed: the same rule applied to a registry that has since gained a row returns a
# different result, and the result alone does not record which set produced it.
#
# A preview run selects its baselines by label and freezes nothing, because a candidate set built
# from reduced results would authorize a selection the reduced run cannot support. It selects by
# label rather than by hash so that the declaration can be written down: a backtest hash is a
# property of the run that produced it, so a preview named by hash can only be launched from the
# machine that has just produced one.
# %%
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
baseline_candidates: CandidateSet | None
if EXECUTION_TIER == "canonical" and (PREVIEW_LABELS or PREVIEW_MAX_BASELINE_CONFIGS):
raise ValueError("canonical execution cannot declare preview reductions")
if EXECUTION_TIER == "canonical":
baseline_population = OfficialPopulation.one(study, name=BASELINE_POPULATION)
baseline_hashes = baseline_population.require_complete()
baseline_table = study.backtests.table().filter(pl.col("backtest_hash").is_in(baseline_hashes))
if baseline_table.height != len(baseline_hashes):
raise RuntimeError("the baseline backtest catalog is incomplete")
baseline_candidates = study.backtests.freeze(
baseline_table,
name=BASELINE_CANDIDATES,
supersedes=candidate_set_supersedes(
study,
name=BASELINE_CANDIDATES,
declared=SUPERSEDES_BASELINE_CANDIDATES or None,
),
)
elif EXECUTION_TIER == "preview":
if not WORKSPACE or not PREVIEW_LABELS or PREVIEW_MAX_BASELINE_CONFIGS < 1:
raise ValueError(
"preview execution requires WORKSPACE, PREVIEW_LABELS and PREVIEW_MAX_BASELINE_CONFIGS"
)
unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
if unknown:
raise ValueError(f"preview labels this case study does not declare: {unknown}")
baseline_table = preview_baseline_candidates(
study, labels=PREVIEW_LABELS, limit=PREVIEW_MAX_BASELINE_CONFIGS
)
baseline_candidates = None
else:
raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
if baseline_table.get_column("sharpe").null_count():
raise RuntimeError("a baseline candidate carries no Sharpe ratio")
# %% [markdown]
# ## Which baselines advance
#
# The shortlist is ordered by validation backtest Sharpe, with the backtest identity breaking
# exact ties so the order does not depend on row order in the registry. Two properties of that
# rule are worth stating, because both are easy to get wrong:
#
# **The unit counted is a model configuration, not a backtest row.** One configuration produced
# several backtests here, one per saved checkpoint and per concentration, and those rows are
# near-duplicates of each other. Each configuration therefore contributes a single row - its
# highest-Sharpe one - and the limit counts distinct configurations, which is what keeps several
# model families on the shortlist.
#
# **Sharpe is the only criterion.** The information coefficient computed upstream measures rank
# correlation between prediction and outcome; it is a diagnostic and selects nothing here, because
# a strategy is chosen on what it earned after costs.
# %%
if TOP_N_PREDICTIONS is None:
TOP_N_PREDICTIONS = get_top_n_predictions(CASE_STUDY, "allocation")
top_n_configs = TOP_N_PREDICTIONS
# A width of 0 asks for every configuration, the spelling `top_n_predictions.signal` uses in
# this setup.yaml. Read as a row count it is `.head(0)`, which selects nothing and leaves the
# count check below comparing 0 against `min(0, available)` - a sweep that registers no
# allocation at all and exits 0.
config_cap = top_n_cap(top_n_configs)
checkpoints_per_config = get_checkpoints_per_config(CASE_STUDY)
ranked = baseline_table.sort("sharpe", "backtest_hash", descending=[True, False])
shortlist = ranked.group_by("family", "config_name", maintain_order=True).head(
checkpoints_per_config
)
if config_cap is not None:
shortlist = shortlist.head(config_cap * checkpoints_per_config)
if baseline_candidates is not None:
best = baseline_candidates.best_validation_sharpe()
if shortlist.item(0, "backtest_hash") != best.hash:
raise RuntimeError("the displayed shortlist disagrees with the candidate-set ranking rule")
available_configs = ranked.select("family", "config_name").n_unique()
expected_configs = available_configs if config_cap is None else min(config_cap, available_configs)
if shortlist.select("family", "config_name").n_unique() != expected_configs:
raise RuntimeError("the allocation shortlist does not hold the declared configuration count")
# %% tags=["results"]
shortlist.select(
"family",
"config_name",
"checkpoint_kind",
"checkpoint_value",
"signal_method",
"sharpe",
"backtest_hash",
)
# %% [markdown]
# ## The weighting rules
#
# Each allocator turns the selected symbols into weights, and the parameters come from
# `config/setup.yaml` so the notebook demonstrates the comparison instead of choosing it:
#
# - **score_weighted** puts capital in proportion to the predicted return, so the model's ranking
# determines position size as well as membership. It concentrates risk exactly where the model
# is most confident, which is what you want if the scores are informative and what hurts most
# if they are not.
# - **inverse_vol** sizes each position by the inverse of its underlying's recent return volatility,
# so a calm name carries more capital than a volatile one.
# - **risk_parity** goes further and solves for weights whose risk contributions are equal, using
# the covariance between underlyings rather than each one's volatility alone.
# - **hrp** clusters the underlyings by how their returns move together and allocates down the
# resulting tree, which avoids inverting a covariance matrix estimated from short samples.
# - **mvo_ledoit_wolf** is mean-variance optimisation with the covariance matrix shrunk toward a
# structured target, the shrinkage being what keeps an estimate from a short window usable.
# - **conformal_weighted** sizes each position by the width of its conformal prediction interval,
# so capital follows how precise the model's forecast is rather than any moment of past returns.
# It is the only rule here that reads the model's own uncertainty, and the only one that trades
# a shorter history than the baseline: an entry date before the first calibration window has no
# prior-only interval to size by, so those cohorts are dropped rather than quietly equal-weighted.
#
# The volatility and covariance windows are all the same length, set once at the case-study level,
# so no allocator is advantaged by seeing more history than another. Equal weight is absent from
# the menu because it is the baseline these are being compared against.
# %%
allocators = get_allocators(CASE_STUDY)
if any(allocation["method"] == "equal_weight" for allocation in allocators):
raise ValueError(
"the allocator menu lists equal_weight, which is the signal stage's own weighting; "
"the comparison would enter the baseline against itself"
)
if EXECUTION_TIER == "preview":
allocators = [row for row in allocators if row["method"] in PREVIEW_ALLOCATORS]
if not allocators:
raise ValueError("allocation request set is empty")
print(f"{len(allocators)} allocators: {sorted(row['method'] for row in allocators)}")
# %% [markdown]
# ## The requests
#
# One request per shortlisted baseline and allocator. Each copies its baseline's signal verbatim -
# the same prediction set, the same concentration, the same liquid universe - and adds the
# allocation block. The only field that differs between a request and the baseline it came from is
# the weighting rule, which is what makes the later comparison a paired one.
# %%
request_rows = []
for row in shortlist.iter_rows(named=True):
baseline = Result.open(
study,
row["backtest_hash"],
include_preview=EXECUTION_TIER == "preview",
)
signal = baseline.spec()["strategy"]["signal"]
for allocation in allocators:
request_rows.append(
{
"request_name": f"{baseline.hash}-{allocation['method']}",
"prediction_hash": row["prediction_hash"],
"label": row["label"],
"baseline_hash": baseline.hash,
"allocation_method": allocation["method"],
"signal": signal,
"allocation": allocation,
"risk": None,
"costs": None,
"chapter": "ch17",
}
)
requests = strategy_request_frame(request_rows)
print(f"{requests.height} requests: {shortlist.height} baselines x {len(allocators)} allocators")
# %% [markdown]
# ## Execute and extend the candidate set
#
# Each request republishes its own decision artifact, because the allocator changes the weights the
# contracts are held at and therefore changes what was traded. The engine then validates the paired
# option lifecycle, that every selected contract ends either by cash settlement or by liquidation,
# the retained hedge, and the cost accounting before the result is published.
#
# The finished results are appended to the frozen baseline set, producing a second named set that
# holds everything selection may consider. Extending creates a new set rather than mutating the old
# one, so the earlier set stays exactly what it was when it was written.
# %%
execution = run_official_backtest_requests(
study,
requests,
population_name=ALLOCATION_POPULATION if EXECUTION_TIER == "canonical" else None,
supersedes=supersedes_for_run(
study,
population_name=ALLOCATION_POPULATION,
declared=SUPERSEDES_ALLOCATION_POPULATION or None,
execution_tier=EXECUTION_TIER,
),
)
catalog = execution.catalog_rows.sort("request_name")
if catalog.height != requests.height or catalog.filter(~pl.col("complete")).height:
raise RuntimeError("allocation execution did not publish every declared request")
strategy_candidates = (
baseline_candidates.extend(
STRATEGY_CANDIDATES,
execution.results,
supersedes=candidate_set_supersedes(
study,
name=STRATEGY_CANDIDATES,
declared=SUPERSEDES_STRATEGY_CANDIDATES or None,
),
)
if baseline_candidates is not None
else None
)
# %% [markdown]
# ## What the run produced
#
# The chart pairs every allocation result against the equal-weight baseline it was built from.
# A point above the diagonal is a baseline the allocator improved on this data; the vertical
# spread within one colour is how much the answer depends on which model the allocator was handed.
# Neither is a selection - that needs the interval around each estimate, which
# `18_strategy_analysis` reports.
#
# Both Sharpe ratios in a pair are recomputed over the dates the two results share, rather than
# read from the registry where each covers its own series. `conformal_weighted` trades a shorter
# history, so its registered number is measured over a different stretch of market than the
# baseline's and the difference between them would carry the period as well as the allocator. The
# summary reports the shortest common support in each row against the length of that same pair's
# baseline, which is how much of the record the thinnest comparison in that row is made on.
# %%
pairs = (
catalog.select("request_name", "backtest_hash")
.join(
requests.select("request_name", "baseline_hash", "allocation_method"),
on="request_name",
how="inner",
)
.join(
baseline_table.select(pl.col("backtest_hash").alias("baseline_hash"), "family"),
on="baseline_hash",
how="inner",
)
)
if pairs.height != catalog.height:
raise RuntimeError("an allocation result did not pair with its baseline")
# Both sides are recomputed on the dates they share. `conformal_weighted` has no weight for an
# entry date with no prior-only calibration window, so it starts trading later than the baseline
# it is built from, and its registered Sharpe covers a different stretch of market.
allocation_sharpe = pairs.join(
paired_sharpe_on_common_support(study, pairs, include_preview=EXECUTION_TIER == "preview"),
on=["backtest_hash", "baseline_hash"],
how="inner",
)
if allocation_sharpe.height != pairs.height:
raise RuntimeError("a pair did not resolve a Sharpe on common support")
# %% tags=["results"]
allocation_summary = (
allocation_sharpe.group_by("allocation_method")
.agg(
backtests=pl.len(),
sharpe_median=pl.col("allocation_sharpe").median(),
improved_on_baseline=(pl.col("allocation_sharpe") > pl.col("baseline_sharpe")).sum(),
# Both from the same pair: baselines within a group differ in length, so a minimum
# overlap taken from one pair and a maximum baseline from another describe no
# comparison in the table.
shortest_common_support=pl.col("n_periods").min(),
its_baseline_sessions=pl.col("baseline_periods").sort_by("n_periods").first(),
)
.sort("allocation_method")
)
allocation_summary
# %%
pairing = px.scatter(
allocation_sharpe,
x="baseline_sharpe",
y="allocation_sharpe",
color="allocation_method",
symbol="family",
hover_data=["baseline_hash", "backtest_hash"],
)
_axis_lo = min(
allocation_sharpe.get_column("baseline_sharpe").min(),
allocation_sharpe.get_column("allocation_sharpe").min(),
)
_axis_hi = max(
allocation_sharpe.get_column("baseline_sharpe").max(),
allocation_sharpe.get_column("allocation_sharpe").max(),
)
pairing.add_shape(
type="line",
x0=_axis_lo,
y0=_axis_lo,
x1=_axis_hi,
y1=_axis_hi,
line=dict(color=COLORS["neutral"], width=1, dash="dash"),
)
pairing.update_layout(
title="Allocated Sharpe against the equal-weight baseline it replaces",
height=560,
width=1000,
margin=dict(t=70),
legend_title_text="allocator",
)
pairing.update_xaxes(title_text="Equal-weight baseline Sharpe")
pairing.update_yaxes(title_text="Allocated Sharpe")
show_plotly_with_alt(
pairing,
"Scatter plot of each allocation backtest's validation Sharpe against the equal-weight "
"baseline it was built from, coloured by allocator, with the diagonal marking no change.",
)
# %% tags=["results"]
pl.DataFrame(
{
"candidate_set": [BASELINE_CANDIDATES, STRATEGY_CANDIDATES],
"member_count": [
len(baseline_candidates.members) if baseline_candidates else 0,
len(strategy_candidates.members) if strategy_candidates else 0,
],
"set_hash": [
baseline_candidates.hash if baseline_candidates else "",
strategy_candidates.hash if strategy_candidates else "",
],
}
)
# %% [markdown]
# ## Key takeaways
#
# - A selection is only reproducible against a recorded candidate set. Freezing the set before
# ranking is what makes "the highest Sharpe" a statement someone else can check.
# - Counting configurations rather than rows is what keeps a shortlist diverse; a checkpoint sweep
# of one model otherwise crowds out every other family without any rule being broken.
# - Holding the signal fixed and varying only the weighting rule is what allows the difference to
# be attributed to the allocator. A comparison that also moved the concentration or the universe
# would confound the three.
#
# **Known limitations**: the covariance-based allocators read the underlying equity's return
# history, not the straddle's, so they size the hedge exposure well and the option exposure only
# indirectly. Every number here is a point estimate over the validation period with no interval
# attached, and the shortlist inherits whatever selection noise the baseline stage carried.
```Exibido na íntegra, com atribuição conforme a licença da fonte. Licença: MIT
Este resumo foi escrito pelo agente de pesquisa da Stratmill com base no original; não é uma cópia da fonte.