עבור לתוכן
כל מסמכי הספרייה

שיטות לקביעת גודל פוזיציה בתיקי חוזים עתידיים של CME

מחברת Machine Learning for Trading

סיכום

המסמך משווה שיטות חלופיות לקביעת גודל פוזיציה באסטרטגיות חוזים עתידיים של CME, שנבחרו מתוך קו בסיס המדורג לפי מדד שארפ באימות עם משקל שווה. שקלול שווה מתייחס לכל מוצר שנבחר באותו אופן, אף שלחוזים עתידיים עשויות להיות רמות תנודתיות ותנועה משותפת שונות מאוד. החלופות משתמשות בתנודתיות של המוצר, בקשרי שונות־משותפת או בקלטים מוגדרים אחרים לקביעת גודל הפוזיציה כדי להתאים חשיפות. שקלול לפי תנודתיות הפוכה דורש אומדן של פחות כמויות ועשוי להיות עמיד יותר כשההיסטוריה מוגבלת, ואילו שיטות המבוססות על שונות־משותפת יכולות להתחשב בפיזור אך חשופות יותר לרעש באומדנים.

תקופות המבט לאחור של מקצה הנכסים קבועות בתצורה ואינן נבחרות לפי מדד שארפ באימות, שכן בחירה כזו הייתה מכווננת את שיטת השקלול לאותם נתונים המשמשים להערכתה. כל אומדן משתמש רק בהיסטוריית מחירים שהייתה זמינה לפני החלטת המסחר. המחברת מעריכה שיטות הקצאה על רשימה קצרה וקבועה של תצורות קו בסיס, ושומרת את התוצאות כאוכלוסייה מוגדרת; שקלול שווה אינו נכלל כי הוא כבר קו הבסיס. הקצאה עשויה לשפר את מדד שארפ באימות ולהישאר זכאית לבחירה, אך ההשוואה היא ברוטו ואינה כוללת את המחזור הנוסף שנגרם משינוי המשקולות. המסמך מפנה לניתוח עלויות נפרד בנושא זה.

רעיונות מרכזיים

  • שקלול שווה עשוי להשאיר את סיכון תיק החוזים העתידיים בשליטת החוזים התנודתיים ביותר.
  • שקלול לפי תנודתיות הפוכה פשוט יותר לאומדן, ואילו שיטות שונות־משותפת מתארות קשרים אך דורשות אומדנים רועשים יותר.
  • קבעו מראש את תקופות המבט לאחור של מקצה הנכסים כדי להימנע מכוונונן לתקופת האימות.
  • חשבו אומדני גודל פוזיציה רק מתוך היסטוריית מחירים הזמינה לפני כל החלטה.
  • התחשבו במחזור הנוסף שנגרם משינוי ההקצאות בעת הערכת ביצועים נטו.

תגיות

הטקסט המלא
# CME Futures: Portfolio Allocation


# CME Futures: Portfolio Allocation

The baseline stage ranks complete configurations by equal-weight validation backtest Sharpe. For
each label, this notebook retains the strongest checkpoint and signal concentration for each of
the configured number of distinct model configurations, then evaluates the declared alternative
position sizing methods. Equal weight is not among them: it is the baseline itself, and because
`stage` is not part of `backtest_hash`, running it again here produces a row hashing
identically to its baseline parent, so one of the two is silently lost. Measured in this
case study's own pre-rebuild store: 48 rows stamped `stage='signal'` while carrying
`allocation.method='equal_weight'`, and no allocation-stage equal-weight rows at all.

All allocator lookbacks come from the case-study configuration. The official population is fixed
before execution; machine speed and caught failures cannot change which allocators run.

## What allocation decides, and why it is not a detail

The baseline held every selected product in the same size. That is a real choice, not the
absence of one, and it says something specific: that the signal's ranking carries information
about *which* products to hold and nothing about *how much* of each.

An allocator disputes that. It sizes positions by some property the ranking does not capture -
how volatile a product has been, how it co-moves with the others, how confident the model was.
The claim is that two products the signal ranks equally are not equally worth the same dollar
risk.

For futures this is a larger effect than it would be in equities, and the reason is in the
instruments. A gold contract and a natural-gas contract with equal notional exposure carry
entirely different risk, because their volatilities differ by a wide margin and they do not
move together. An equal-weight book of futures is therefore not a neutral book: it is one
whose realized risk is dominated by whichever contracts happen to be most volatile, and its
overall behaviour can be driven by two or three positions out of thirty regardless of what the
signal said.

### What the alternatives are actually doing

The declared methods differ in how much they estimate, and that is the axis to read them on.
Sizing inversely to a product's own volatility uses one number per product and no relationship
between them - it equalizes risk contribution under the assumption that the products are
independent, which they are not, but it is robust because a single volatility is estimated
accurately from little data. Methods that use the covariance between products can in principle
do better, because diversification is a property of the relationships rather than of any one
series. In practice they must estimate far more quantities from the same history, and a
covariance matrix estimated from a short window is dominated by noise that the optimizer then
treats as signal - which is why the more sophisticated method is not reliably the better one
and why they are compared here rather than assumed.

### Why the lookbacks are declared in configuration

Every allocator reads a history to estimate from, and the length of that history is a free
parameter that materially changes the result: a short window tracks a regime change quickly
and is noisy, a long one is stable and stale. Choosing it by validation Sharpe would be fitting
the allocator to the same path it is then assessed on, and the improvement would be
indistinguishable from a real one. The lookbacks come from `config/setup.yaml` for the same
reason the risk-overlay parameters do.

### Why equal weight is excluded here rather than re-run

Equal weight is the baseline, and the paragraph above the parameter cell records what happens
if it is run again in this stage: because `stage` is not part of `backtest_hash`, the row
hashes identically to its baseline parent and one of the two is silently lost. Measured in
this case study's own pre-rebuild store, 48 rows carried `stage='signal'` while declaring
`allocation.method='equal_weight'`, and there were no allocation-stage equal-weight rows at
all. The comparison a reader wants - allocator against equal weight - is made against the
baseline population, not by recomputing it here.

### These candidates stay eligible

Allocation results are part of the final selection pool, so an allocator that genuinely
improves validation Sharpe can be what the case study ships. One consequence to carry forward:
every allocator except equal weight re-sizes positions as its estimates move, which is
turnover the baseline did not have. `16_costs` is where that is priced, and an allocator's
advantage here is a gross number.

```python
"""Run the declared CME futures allocation population."""

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    create_label_candidate_sets,
    open_study,
    product_universe_table,
    run_official_backtest_requests,
    shortlist_signal_configurations,
    strategy_request_frame,
)
from case_studies.research.population import supersedes_for_run
from case_studies.utils.sweep_config import get_allocators, get_top_n_predictions
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []
PREVIEW_MAX_BASELINE_ROWS = 0
# None means the width `setup.yaml` declares; an int overrides it. Declared here because
# papermill only binds a name the parameters cell already holds - a run that passes
# TOP_N_PREDICTIONS to a notebook without it sweeps the declared width and exits 0.
TOP_N_PREDICTIONS = None

# The allocation population is immutable under its name, so a run whose members have moved has
# to say which generation it retires. Anything upstream that changes a backtest identity moves
# them - a corrected label, a changed accounting field, a re-run after a registry reset - and
# `OfficialPopulation.create` refuses to write a different member list under a name that already
# exists. Declared as a literal so that running the committed notebook as it stands recomputes
# the population on record. Empty for a first snapshot.
ALLOCATION_POPULATION = "cme_futures-allocation-validation-v1"
SUPERSEDES_ALLOCATION_POPULATION: str = ""

# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "cme_futures-allocation-fwd_ret_5d-v1": "live",
    "cme_futures-allocation-fwd_ret_21d-v1": "live",
}
```

## Select signal configurations by validation Sharpe

The shortlist is deterministic. It scans the immutable signal candidate set in descending Sharpe
order with the backtest identity as tie-break, and keeps one exact checkpoint and strategy per
distinct `(family, config_name)` pair.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
    if PREVIEW_LABELS or PREVIEW_MAX_BASELINE_ROWS:
        raise ValueError("canonical execution cannot declare preview reductions")
    labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
    if WORKSPACE is None or not PREVIEW_LABELS or PREVIEW_MAX_BASELINE_ROWS < 1:
        raise ValueError(
            "preview execution requires WORKSPACE, PREVIEW_LABELS and PREVIEW_MAX_BASELINE_ROWS"
        )
    unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
    if unknown:
        raise ValueError(f"preview labels this case study does not declare: {unknown}")
    labels = tuple(PREVIEW_LABELS)
else:
    raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
universe = product_universe_table()
universe
```

## How many baseline configurations the position sizing methods run on

**Why a shortlist rather than every baseline row.** Allocation is applied to the strongest
configurations rather than to all of them, and the reason is that the question is about the
allocator, not about the configuration underneath it. Running every allocator against every
baseline row would multiply the population by the number of methods and answer a question
nobody asked, while making the effective number of trials far larger - which the deflated
Sharpe downstream then has to divide by. Keeping the strongest checkpoint and concentration
per configuration asks the allocator's question against the strategies that would otherwise
have shipped.

Canonical takes the shortlist size from `setup.yaml`, which is the declared width of the
allocation stage. A preview cannot: it backtests a bounded slice of the baseline stage, so the
canonical width names more distinct configurations than its pool contains and
`shortlist_signal_configurations` refuses - correctly, since silently returning fewer is the
quiet shrinking that strictness exists to prevent. The preview therefore declares its own
width, and is held to it just as strictly.

`TOP_N_PREDICTIONS` overrides both, which is how a sweep runs wider than the shipped
declaration. It is held to its width the same way: a number the pool cannot fill raises here
rather than quietly shrinking. A sweep that wants every configuration passes `0`, the
spelling `top_n_predictions.signal` uses one line above the declaration this overrides. That
is not the same request as a number large enough to be sure: 999 against a population of 50
is indistinguishable from a population that shrank to 50, and stays a refusal.

```python
if TOP_N_PREDICTIONS is None:
    TOP_N_PREDICTIONS = (
        get_top_n_predictions("cme_futures", "allocation")
        if EXECUTION_TIER == "canonical"
        else PREVIEW_MAX_BASELINE_ROWS
    )
shortlist_size = TOP_N_PREDICTIONS
allocators = get_allocators("cme_futures")
if not allocators:
    raise ValueError("the configured allocator population is empty")
if any(allocation.get("method") == "equal_weight" for allocation in allocators):
    raise ValueError(
        "equal_weight is the baseline stage, not an allocator: `stage` is not part of "
        "`backtest_hash`, so an equal-weight reweight hashes identically to its baseline "
        "parent and one of the two rows is lost. Remove it from the configured menu."
    )

request_rows = []
for label in labels:
    for baseline in shortlist_signal_configurations(
        study,
        label=label,
        limit=shortlist_size,
        execution_tier=EXECUTION_TIER,
    ):
        prediction_hash = baseline.registry_record()["prediction_hash"]
        signal = baseline.spec()["strategy"]["signal"]
        for allocation in allocators:
            method = allocation["method"]
            request_rows.append(
                {
                    "request_name": f"{baseline.hash}-{method}",
                    "prediction_hash": prediction_hash,
                    "label": label,
                    "signal": signal,
                    "allocation": allocation,
                    "risk": None,
                    "costs": None,
                    "chapter": "ch17",
                }
            )
requests = strategy_request_frame(request_rows)
requests.select("request_name", "prediction_hash", "label", "signal", "allocation")
```

## Execute and freeze allocation candidates

Moment-based allocators receive only price history before each decision. Product-keyed typed
decisions retain the selected prediction, roll audit, expiry reference, and allocation settings.

"Only price history before each decision" is the causality condition for this stage, and it is
the one an allocator makes easy to violate: a covariance or volatility estimate computed over
the whole sample would size every position using dispersion the market had not yet shown, and
the resulting book would look well-balanced for reasons unavailable at the time. The estimate
is rebuilt at each decision from history strictly before it.

The results freeze as a named population, and `SUPERSEDES_ALLOCATION_POPULATION` in the
parameter cell names the generation a re-run retires. The retired snapshot stays in the
registry, so a Sharpe quoted from it remains traceable to the set it was computed over.

```python
execution = run_official_backtest_requests(
    study,
    requests,
    population_name=ALLOCATION_POPULATION if EXECUTION_TIER == "canonical" else None,
    supersedes=supersedes_for_run(
        study,
        population_name=ALLOCATION_POPULATION,
        declared=SUPERSEDES_ALLOCATION_POPULATION or None,
        execution_tier=EXECUTION_TIER,
    ),
)
candidate_sets = (
    create_label_candidate_sets(
        study, execution, stage="allocation", supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
    )
    if EXECUTION_TIER == "canonical"
    else {}
)
```

`source` says whether each member was computed by this run or served from the registry because
an identical identity was already recorded. A re-run of a registered sweep is entirely `reused`
and completes in seconds; without the column that is indistinguishable from having computed
every row.

```python
execution.catalog_rows.sort("label", "request_name")
```

The next two execution notebooks select the highest validation Sharpe from the union of signal and
allocation results for each label. Cost sensitivity is diagnostic; risk overlays remain eligible
for final selection.

מוצג במלואו בציון המקור ובהתאם לרישיון שלו. רישיון: MIT

הסיכום נכתב בידי סוכן המחקר של Stratmill על סמך המקור; הוא אינו העתק של המקור.