コンテンツへスキップ
ライブラリの全資料

CME先物ポートフォリオのポジションサイズ手法

ノートブック Machine Learning for Trading

サマリー

この文書では、均等加重の検証シャープレシオで順位付けしたベースラインから選んだCME先物戦略について、複数のポジションサイズ手法を比較します。均等加重では選択した商品をすべて同じように扱いますが、先物契約によってボラティリティや連動性は大きく異なる場合があります。代替手法では、商品のボラティリティ、共分散関係、その他明示したサイズ設定の入力値を使ってエクスポージャーを調整します。逆ボラティリティによるサイズ設定は推定項目が少なく、履歴が限られる場合に頑健性が高い可能性があります。一方、共分散に基づく手法は分散効果を考慮できますが、推定ノイズの影響を受けやすくなります。

アロケーターのルックバック期間は、評価に使うデータに合わせて調整される検証シャープレシオから選ばず、設定で固定します。各推定には、取引判断前に利用可能だった価格履歴だけを使います。ノートブックでは、ベースラインの固定された候補リストに対して配分手法を評価し、結果を名前付きの集合として保存します。均等加重はすでにベースラインなので除外します。配分変更で検証シャープレシオが改善すれば選択対象に残せますが、この比較にはウェイト変更による追加売買回転のコストが含まれていません。その影響については、別のコスト分析を参照するよう案内しています。

主なアイデア

  • 均等加重では、ボラティリティが最も高い契約が先物ポートフォリオのリスクを支配する場合があります。
  • 逆ボラティリティによるサイズ設定は推定が比較的単純です。共分散手法は相互関係をモデル化しますが、よりノイズの大きい推定が必要です。
  • 検証期間に合わせた調整を避けるため、アロケーターのルックバック期間は事前に固定します。
  • サイズの推定には、各判断の前に利用可能だった価格履歴だけを使います。
  • ネットパフォーマンスを評価する際は、配分変更による追加の売買回転を考慮します。

タグ

全文
# CME Futures: Portfolio Allocation


# CME Futures: Portfolio Allocation

The baseline stage ranks complete configurations by equal-weight validation backtest Sharpe. For
each label, this notebook retains the strongest checkpoint and signal concentration for each of
the configured number of distinct model configurations, then evaluates the declared alternative
position sizing methods. Equal weight is not among them: it is the baseline itself, and because
`stage` is not part of `backtest_hash`, running it again here produces a row hashing
identically to its baseline parent, so one of the two is silently lost. Measured in this
case study's own pre-rebuild store: 48 rows stamped `stage='signal'` while carrying
`allocation.method='equal_weight'`, and no allocation-stage equal-weight rows at all.

All allocator lookbacks come from the case-study configuration. The official population is fixed
before execution; machine speed and caught failures cannot change which allocators run.

## What allocation decides, and why it is not a detail

The baseline held every selected product in the same size. That is a real choice, not the
absence of one, and it says something specific: that the signal's ranking carries information
about *which* products to hold and nothing about *how much* of each.

An allocator disputes that. It sizes positions by some property the ranking does not capture -
how volatile a product has been, how it co-moves with the others, how confident the model was.
The claim is that two products the signal ranks equally are not equally worth the same dollar
risk.

For futures this is a larger effect than it would be in equities, and the reason is in the
instruments. A gold contract and a natural-gas contract with equal notional exposure carry
entirely different risk, because their volatilities differ by a wide margin and they do not
move together. An equal-weight book of futures is therefore not a neutral book: it is one
whose realized risk is dominated by whichever contracts happen to be most volatile, and its
overall behaviour can be driven by two or three positions out of thirty regardless of what the
signal said.

### What the alternatives are actually doing

The declared methods differ in how much they estimate, and that is the axis to read them on.
Sizing inversely to a product's own volatility uses one number per product and no relationship
between them - it equalizes risk contribution under the assumption that the products are
independent, which they are not, but it is robust because a single volatility is estimated
accurately from little data. Methods that use the covariance between products can in principle
do better, because diversification is a property of the relationships rather than of any one
series. In practice they must estimate far more quantities from the same history, and a
covariance matrix estimated from a short window is dominated by noise that the optimizer then
treats as signal - which is why the more sophisticated method is not reliably the better one
and why they are compared here rather than assumed.

### Why the lookbacks are declared in configuration

Every allocator reads a history to estimate from, and the length of that history is a free
parameter that materially changes the result: a short window tracks a regime change quickly
and is noisy, a long one is stable and stale. Choosing it by validation Sharpe would be fitting
the allocator to the same path it is then assessed on, and the improvement would be
indistinguishable from a real one. The lookbacks come from `config/setup.yaml` for the same
reason the risk-overlay parameters do.

### Why equal weight is excluded here rather than re-run

Equal weight is the baseline, and the paragraph above the parameter cell records what happens
if it is run again in this stage: because `stage` is not part of `backtest_hash`, the row
hashes identically to its baseline parent and one of the two is silently lost. Measured in
this case study's own pre-rebuild store, 48 rows carried `stage='signal'` while declaring
`allocation.method='equal_weight'`, and there were no allocation-stage equal-weight rows at
all. The comparison a reader wants - allocator against equal weight - is made against the
baseline population, not by recomputing it here.

### These candidates stay eligible

Allocation results are part of the final selection pool, so an allocator that genuinely
improves validation Sharpe can be what the case study ships. One consequence to carry forward:
every allocator except equal weight re-sizes positions as its estimates move, which is
turnover the baseline did not have. `16_costs` is where that is priced, and an allocator's
advantage here is a gross number.

```python
"""Run the declared CME futures allocation population."""

from case_studies.cme_futures.research_workflow import (
    ALL_LABELS,
    create_label_candidate_sets,
    open_study,
    product_universe_table,
    run_official_backtest_requests,
    shortlist_signal_configurations,
    strategy_request_frame,
)
from case_studies.research.population import supersedes_for_run
from case_studies.utils.sweep_config import get_allocators, get_top_n_predictions
```

```python
EXECUTION_TIER = "canonical"
WORKSPACE: str | None = None
PREVIEW_LABELS: list[str] = []
PREVIEW_MAX_BASELINE_ROWS = 0
# None means the width `setup.yaml` declares; an int overrides it. Declared here because
# papermill only binds a name the parameters cell already holds - a run that passes
# TOP_N_PREDICTIONS to a notebook without it sweeps the declared width and exits 0.
TOP_N_PREDICTIONS = None

# The allocation population is immutable under its name, so a run whose members have moved has
# to say which generation it retires. Anything upstream that changes a backtest identity moves
# them - a corrected label, a changed accounting field, a re-run after a registry reset - and
# `OfficialPopulation.create` refuses to write a different member list under a name that already
# exists. Declared as a literal so that running the committed notebook as it stands recomputes
# the population on record. Empty for a first snapshot.
ALLOCATION_POPULATION = "cme_futures-allocation-validation-v1"
SUPERSEDES_ALLOCATION_POPULATION: str = ""

# The per-label candidate sets this notebook freezes are immutable under their names too, and
# for the same reason as the population above: `CandidateSet.create` refuses a changed member
# list under a name that already exists. Nothing reached that argument before, so any run whose
# membership moved - which a wider sweep does by construction - stopped at the freeze after the
# fit, with no parameter able to answer it.
#
# Each name maps to the generation this run retires. `"live"` names the lineage and looks the
# generation up, which is the form that does not decay: naming the head instead is correct only
# until the next publish, because `create` accepts the head and nothing else. The declaration is
# resolved through `candidate_set_supersedes` rather than offered straight, so a reader's clean
# clone - which has no generation to replace, and often no `candidate_sets` table at all -
# publishes generation one instead of being refused. An unchanged re-run never reads it: a set's
# hash is computed from its members and its contract, so the existing name binding answers.
SUPERSEDES_CANDIDATE_SETS: dict[str, str] = {
    "cme_futures-allocation-fwd_ret_5d-v1": "live",
    "cme_futures-allocation-fwd_ret_21d-v1": "live",
}
```

## Select signal configurations by validation Sharpe

The shortlist is deterministic. It scans the immutable signal candidate set in descending Sharpe
order with the backtest identity as tie-break, and keeps one exact checkpoint and strategy per
distinct `(family, config_name)` pair.

```python
study = open_study(execution_tier=EXECUTION_TIER, workspace=WORKSPACE)
if EXECUTION_TIER == "canonical":
    if PREVIEW_LABELS or PREVIEW_MAX_BASELINE_ROWS:
        raise ValueError("canonical execution cannot declare preview reductions")
    labels = ALL_LABELS
elif EXECUTION_TIER == "preview":
    if WORKSPACE is None or not PREVIEW_LABELS or PREVIEW_MAX_BASELINE_ROWS < 1:
        raise ValueError(
            "preview execution requires WORKSPACE, PREVIEW_LABELS and PREVIEW_MAX_BASELINE_ROWS"
        )
    unknown = sorted(set(PREVIEW_LABELS) - set(ALL_LABELS))
    if unknown:
        raise ValueError(f"preview labels this case study does not declare: {unknown}")
    labels = tuple(PREVIEW_LABELS)
else:
    raise ValueError(f"unsupported execution tier: {EXECUTION_TIER!r}")
universe = product_universe_table()
universe
```

## How many baseline configurations the position sizing methods run on

**Why a shortlist rather than every baseline row.** Allocation is applied to the strongest
configurations rather than to all of them, and the reason is that the question is about the
allocator, not about the configuration underneath it. Running every allocator against every
baseline row would multiply the population by the number of methods and answer a question
nobody asked, while making the effective number of trials far larger - which the deflated
Sharpe downstream then has to divide by. Keeping the strongest checkpoint and concentration
per configuration asks the allocator's question against the strategies that would otherwise
have shipped.

Canonical takes the shortlist size from `setup.yaml`, which is the declared width of the
allocation stage. A preview cannot: it backtests a bounded slice of the baseline stage, so the
canonical width names more distinct configurations than its pool contains and
`shortlist_signal_configurations` refuses - correctly, since silently returning fewer is the
quiet shrinking that strictness exists to prevent. The preview therefore declares its own
width, and is held to it just as strictly.

`TOP_N_PREDICTIONS` overrides both, which is how a sweep runs wider than the shipped
declaration. It is held to its width the same way: a number the pool cannot fill raises here
rather than quietly shrinking. A sweep that wants every configuration passes `0`, the
spelling `top_n_predictions.signal` uses one line above the declaration this overrides. That
is not the same request as a number large enough to be sure: 999 against a population of 50
is indistinguishable from a population that shrank to 50, and stays a refusal.

```python
if TOP_N_PREDICTIONS is None:
    TOP_N_PREDICTIONS = (
        get_top_n_predictions("cme_futures", "allocation")
        if EXECUTION_TIER == "canonical"
        else PREVIEW_MAX_BASELINE_ROWS
    )
shortlist_size = TOP_N_PREDICTIONS
allocators = get_allocators("cme_futures")
if not allocators:
    raise ValueError("the configured allocator population is empty")
if any(allocation.get("method") == "equal_weight" for allocation in allocators):
    raise ValueError(
        "equal_weight is the baseline stage, not an allocator: `stage` is not part of "
        "`backtest_hash`, so an equal-weight reweight hashes identically to its baseline "
        "parent and one of the two rows is lost. Remove it from the configured menu."
    )

request_rows = []
for label in labels:
    for baseline in shortlist_signal_configurations(
        study,
        label=label,
        limit=shortlist_size,
        execution_tier=EXECUTION_TIER,
    ):
        prediction_hash = baseline.registry_record()["prediction_hash"]
        signal = baseline.spec()["strategy"]["signal"]
        for allocation in allocators:
            method = allocation["method"]
            request_rows.append(
                {
                    "request_name": f"{baseline.hash}-{method}",
                    "prediction_hash": prediction_hash,
                    "label": label,
                    "signal": signal,
                    "allocation": allocation,
                    "risk": None,
                    "costs": None,
                    "chapter": "ch17",
                }
            )
requests = strategy_request_frame(request_rows)
requests.select("request_name", "prediction_hash", "label", "signal", "allocation")
```

## Execute and freeze allocation candidates

Moment-based allocators receive only price history before each decision. Product-keyed typed
decisions retain the selected prediction, roll audit, expiry reference, and allocation settings.

"Only price history before each decision" is the causality condition for this stage, and it is
the one an allocator makes easy to violate: a covariance or volatility estimate computed over
the whole sample would size every position using dispersion the market had not yet shown, and
the resulting book would look well-balanced for reasons unavailable at the time. The estimate
is rebuilt at each decision from history strictly before it.

The results freeze as a named population, and `SUPERSEDES_ALLOCATION_POPULATION` in the
parameter cell names the generation a re-run retires. The retired snapshot stays in the
registry, so a Sharpe quoted from it remains traceable to the set it was computed over.

```python
execution = run_official_backtest_requests(
    study,
    requests,
    population_name=ALLOCATION_POPULATION if EXECUTION_TIER == "canonical" else None,
    supersedes=supersedes_for_run(
        study,
        population_name=ALLOCATION_POPULATION,
        declared=SUPERSEDES_ALLOCATION_POPULATION or None,
        execution_tier=EXECUTION_TIER,
    ),
)
candidate_sets = (
    create_label_candidate_sets(
        study, execution, stage="allocation", supersedes_by_set=SUPERSEDES_CANDIDATE_SETS
    )
    if EXECUTION_TIER == "canonical"
    else {}
)
```

`source` says whether each member was computed by this run or served from the registry because
an identical identity was already recorded. A re-run of a registered sweep is entirely `reused`
and completes in seconds; without the column that is indistinguishable from having computed
every row.

```python
execution.catalog_rows.sort("label", "request_name")
```

The next two execution notebooks select the highest validation Sharpe from the union of signal and
allocation results for each label. Cost sensitivity is diagnostic; risk overlays remain eligible
for final selection.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT

この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。