FX 페어 순위 모델의 동일 비중 기준선 구성
코드 Machine Learning for Trading
요약
이 노트북은 통화 페어 예측 순위를 실제 매매 기준선으로 전환합니다. 각 의사결정 시점에 순위가 가장 높고 낮은 페어를 같은 규모의 롱 및 숏 포지션군으로 구성한 다음 FX 백테스트 엔진으로 평가합니다. 동일 비중을 사용해 포지션 규모와 배분에 관한 후속 결정과 순위의 기여도를 분리합니다. 포트폴리오 규모 그리드는 포지션군의 집중도를 바꾸며, 유니버스 규모는 양쪽에 보유할 수 있는 서로 다른 페어 수를 제한합니다.
노트북은 완전한 검증 예측 모집단을 고정하고 이후 모델 세대에서 퇴출된 식별자를 제외하며, 전체 실행 목록을 계획하고 계획된 모든 백테스트가 완료됐는지 확인합니다. 순위 상관계수나 AUC 같은 예측 통계가 실제 매매 수익을 보장하지는 않는다고 설명합니다. 유동성과 선택한 순위 구간이 결과에 영향을 줄 수 있습니다. 여러 설정의 검증 샤프는 선택 과정으로 부풀려질 수도 있으므로, 이 노트북은 기준선을 비교점으로 삼고 최종 평가를 위해 봉인된 홀드아웃을 남겨 둡니다. 결과는 단 하나의 과거 검증 표본에 한정됩니다.
핵심 아이디어
- 동일 비중 롱 및 숏 포지션군을 사용하면 순위의 질과 배분 선택을 분리해 볼 수 있습니다.
- 예측 순위를 트레이더가 경험하는 수익으로 전환하려면 포지션으로 구성해야 합니다.
- 모집단을 완전하게 추적하면 실행 누락이나 실패가 스윕에서 드러납니다.
- 현재 기준선 모집단을 정할 때 퇴출된 예측 세대는 제외해야 합니다.
- 설정 그리드가 크면 검증 성과가 낙관적인 우승 결과를 낼 수 있으므로 최종 평가에는 홀드아웃이 필요합니다.
태그
전문
# 13_backtest.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Equal-Weight Backtest - FX Pairs
#
# This notebook establishes the strategy baseline. At each decision time it ranks the currency-pair
# scores, takes equal-sized long and short sleeves, and runs the existing FX backtest engine. Every
# complete model configuration and checkpoint is evaluated. The input is a Polars catalog selection,
# so no prediction hash is copied into orchestration code.
#
# This is the step that turns a model into a strategy. Everything before it scored pairs; nothing
# before it has said what you would have held or what it would have returned. Rank correlation
# and AUC measure whether the ordering carries information, and a model can be clearly informative
# by both and still lose money once the ordering is turned into positions - the information may
# sit in pairs too illiquid to hold, or in a part of the distribution the sleeves never reach.
# The return series produced here is the first quantity in the chain that a trader could have
# experienced, and it is the reference every later stage is measured against.
#
# Equal weight is used deliberately rather than for convenience. It is the choice not to choose a
# size, so the resulting return series carries the ranking's contribution and nothing else. Any
# other weighting would blend the model's information with a view about sizing, and the
# allocation notebook could then no longer attribute its improvement to allocation.
#
# **A note on how this appears in the registry.** These runs register with `stage='signal'`, and
# the notebook asserts it. That stored value names the strategy component being varied - here the
# signal mapping - and it does not mean there is a separate "signal stage" upstream of a backtest.
# These *are* the baseline backtests. In prose, write "the baseline backtests" and "baseline
# (equal-weight) Sharpe", never "signal Sharpe", which reads as a metric of the predictions rather
# than of a traded strategy.
#
# **Learning objectives**
#
# - Freeze the exact prediction population before running a baseline sweep.
# - Send selected catalog rows to the canonical FX engine.
# - Verify that every declared prediction and portfolio-size combination produces one backtest.
# - Read the baseline as the reference the allocation, risk and cost stages are measured against.
#
# **Book reference**: Chapter 16
#
# **Prerequisite**: `12_model_analysis`.
# %%
"""Run the complete equal-weight FX validation backtest population."""
import polars as pl
import yaml
from case_studies.research import (
OfficialPopulation,
open_study,
plan_backtests,
population_supersedes,
research_name,
reuse_disclosure,
run_backtests,
superseded_members,
)
from case_studies.utils.sweep_config import get_top_k_values_for
from utils.paths import get_case_study_dir
from utils.reproducibility import set_global_seeds
# %% tags=["parameters"]
CASE_STUDY_ID = "fx_pairs"
EXECUTION_TIER = "canonical"
WORKSPACE: str = ""
LABEL = ""
SPLIT = "validation"
TOP_K = 0
SEED = 42
RUN_SWEEP = True
FORCE_REBACKTEST = False
TOP_N_PREDICTIONS = None
POPULATION_NAME = ""
# Both name the lineage rather than a generation of it. The hashes these replace -
# `59edb30b6271` and `16787d96bf75` - are in no lineage the registry holds, and the names
# they publish under are built by `research_name`, so nothing could look them up to say
# whether they were dead or waiting for a first publication. `"live"` does not have to
# decide: it resolves against the name this run publishes under, which is the one thing
# both call sites already know. See `case_studies.research.population.SUPERSEDES_LIVE`.
SUPERSEDES_EQUAL_WEIGHT_BASELINES: str = "live"
SUPERSEDES_VALIDATION_PREDICTIONS: str = "live"
# %% [markdown]
# ## Select the exact prediction population
#
# Canonical production includes every complete validation prediction the model notebooks currently
# publish. Label, configuration, or population limits are preview controls and cannot define the
# official baseline.
#
# **"Currently publish" is a question about lineage, and the catalog cannot answer it.** A row's
# `identity_status` is derived from the schema version it was written under, so it says the registry
# still understands the row - not that the row is the one its producer stands behind. The two agree
# until a model notebook refits. Then it publishes a second generation of its population under the
# same name, the first generation's prediction sets stay in the registry complete and current, and a
# sweep selecting on the catalog alone runs over both. It would not fail; it would report every
# member complete, over twice the population, and freeze the retired half into the baseline the rest
# of the case study is measured against.
#
# `superseded_members` asks the registry which identities a later generation retired and no
# generation in force still lists, which is exactly the set to drop. The `tabular_dl` and
# `deep_learning` populations each have a retired generation here, because a training identity
# covers the runner's own source file and both runners changed after their first fits were
# registered. The count of what that excludes is printed rather than left implicit.
#
# **A published population can need a second generation too.** The two names this notebook
# publishes are lists of identities, and the exclusion above changes both of them the moment a
# model notebook refits: the prediction population loses the retired members, and every baseline
# backtest resolved from them goes with it. `OfficialPopulation.create` refuses a changed list
# under an existing name without being told which snapshot it replaces, so each name has its own
# declaration and each is offered through `population_supersedes` on the same rule the model
# notebooks use. Both are empty here because neither name has published a first generation yet;
# after the first canonical run, a refit upstream is answered by filling in the hash that run
# printed.
# %% tags=["results"]
set_global_seeds(SEED)
if SPLIT != "validation":
raise ValueError("the baseline sweep uses validation predictions")
if FORCE_REBACKTEST:
raise ValueError("identical complete backtests are reused by identity")
if not RUN_SWEEP:
raise ValueError("set RUN_SWEEP=True to execute the visible baseline request")
study = open_study(CASE_STUDY_ID, execution_tier=EXECUTION_TIER, workspace=WORKSPACE or None)
# The execution tier decides which registry namespace this run reads and writes;
# the reduction knobs decide only how much of it is covered. Inferring the tier
# from the knobs conflated the two, so any reduced run went looking for preview
# predictions - and a reduced run over a canonical upstream, which is what the
# test suite exercises, then resolved no rows at all.
include_preview = EXECUTION_TIER == "preview"
catalog = study.predictions.table(include_preview=include_preview).filter(
(pl.col("identity_status") == "current") & (pl.col("split") == SPLIT) & pl.col("complete")
)
if include_preview:
catalog = catalog.filter(pl.col("execution_tier") == "preview")
else:
catalog = catalog.filter(pl.col("execution_tier") == "canonical")
retired = superseded_members(study, member_kind="prediction")
if retired:
offered = catalog.height
catalog = catalog.filter(~pl.col("prediction_hash").is_in(list(retired)))
print(f"Retired by a later generation, excluded: {offered - catalog.height} of {offered}")
if LABEL:
catalog = catalog.filter(pl.col("label") == LABEL)
if TOP_N_PREDICTIONS is not None:
catalog = catalog.sort("label", "family", "config_name", "checkpoint_value").head(
TOP_N_PREDICTIONS
)
if catalog.is_empty():
raise RuntimeError("the baseline request resolved no complete prediction rows")
# TOP_K and TOP_N_PREDICTIONS both narrow what is backtested, and a narrowed run declares
# a different set of members than the canonical population does. A population is immutable
# once written, so such a run must publish under its own name rather than register a
# partial snapshot under the canonical one. The tier is a separate question: a canonical
# run may legitimately be narrowed, it just may not claim to be the whole population.
if (
(TOP_K or TOP_N_PREDICTIONS is not None or LABEL)
and not include_preview
and not POPULATION_NAME
):
raise ValueError(
"this run narrows the baseline sweep, so it cannot publish the canonical "
"population; pass POPULATION_NAME to give it its own"
)
if catalog.get_column("prediction_hash").n_unique() != catalog.height:
raise RuntimeError("the baseline population contains duplicate prediction identities")
if not include_preview:
predictions_name = research_name(CASE_STUDY_ID, "validation-predictions", scope=POPULATION_NAME)
prediction_population = OfficialPopulation.create(
study,
name=predictions_name,
member_kind="prediction",
members=catalog.get_column("prediction_hash").to_list(),
supersedes=population_supersedes(
study, name=predictions_name, declared=SUPERSEDES_VALIDATION_PREDICTIONS
),
)
prediction_population.require_complete()
print(f"Frozen prediction population: {prediction_population.hash}")
else:
print("Preview selection is isolated from the official prediction population.")
catalog.select(
"label",
"family",
"config_name",
"checkpoint_kind",
"checkpoint_value",
"prediction_hash",
)
# %% [markdown]
# ## Build the baseline strategy grid
#
# `top_k` is the number of pairs in each sleeve. This account allows short selling, so the engine
# holds `top_k` long and `top_k` short and clamps each sleeve to half the universe to keep them
# disjoint. Two consequences worth stating exactly, because both shape how the grid reads:
#
# - At `top_k` below half the universe, ranking decides *which* pairs are held.
# - At `top_k` equal to half the universe, every pair is held at every rebalance and ranking decides
# only the *side*. The grid then varies concentration and sign exposure, not membership.
#
# A `top_k` above half the universe cannot express anything the half-universe value does not: the
# engine clamps it to the same sleeves, so it would register a distinct identity carrying an
# identical weight series. The grid is refused rather than allowed to contain that duplicate.
# %%
universe_symbols = yaml.safe_load(
(get_case_study_dir(CASE_STUDY_ID) / "config" / "setup.yaml").read_text()
)["universe"]["symbols"]
n_assets = len(universe_symbols)
max_sleeve = n_assets // 2
jobs = []
for label in sorted(catalog.get_column("label").unique()):
selected = catalog.filter(pl.col("label") == label)
top_k_values = [TOP_K] if TOP_K else get_top_k_values_for(CASE_STUDY_ID, label, n_assets)
duplicates = sorted({value for value in top_k_values if value > max_sleeve})
if duplicates:
raise RuntimeError(
f"top_k {duplicates} exceed the {max_sleeve}-pair sleeve ceiling for a long-short "
f"account on {n_assets} pairs; the engine would clamp them to {max_sleeve} and register "
"a duplicate weight series under a distinct identity"
)
for top_k in top_k_values:
jobs.append(
{
"label": label,
"top_k": top_k,
"predictions": selected,
"expected": selected.height,
}
)
pl.DataFrame(
[
{"label": job["label"], "top_k": job["top_k"], "prediction_sets": job["expected"]}
for job in jobs
]
)
# %% [markdown]
# ## Freeze the expected baseline population
#
# Planning resolves every backtest identity without running or writing it. Production freezes that
# complete expected set before the first member executes, so a failed member remains visible.
#
# Visibility is the whole point, and it is worth being explicit about the alternative. A sweep that
# collected whatever finished would publish a baseline whose membership depends on which runs
# happened to succeed. Failures are not random with respect to the thing being measured: a
# configuration that produces degenerate scores, or one whose predictions are too sparse to fill a
# sleeve, is more likely to fail and would silently leave the population. The surviving baseline
# would then look better than the sweep that produced it, and nothing in the output would record
# that anything was dropped. Declaring the expected set first converts that from a silent
# improvement into an incomplete population that refuses to publish.
# %% tags=["results"]
planned_hashes = []
for job in jobs:
plan = plan_backtests(
study,
predictions=job["predictions"],
signal={"method": "equal_weight_top_k", "top_k": job["top_k"]},
chapter="16",
)
if len(plan.members) != job["expected"]:
raise RuntimeError("a baseline plan omitted a selected prediction")
planned_hashes.extend(plan.expected_hashes)
if len(planned_hashes) != len(set(planned_hashes)):
raise RuntimeError("two planned baseline jobs collapse to the same backtest identity")
baseline_population = None
if not include_preview:
baselines_name = research_name(CASE_STUDY_ID, "equal-weight-baselines", scope=POPULATION_NAME)
baseline_population = OfficialPopulation.create(
study,
name=baselines_name,
member_kind="backtest",
members=planned_hashes,
supersedes=population_supersedes(
study, name=baselines_name, declared=SUPERSEDES_EQUAL_WEIGHT_BASELINES
),
)
print(f"Frozen expected baseline population: {baseline_population.hash}")
else:
print("Preview backtests remain outside official populations and selection.")
# %% [markdown]
# ## Run every catalog row through the FX engine, then validate what was frozen
#
# Each selected row produces an independent backtest. The loop has no exception-and-continue path:
# one failed member leaves the predeclared population incomplete and stops publication.
#
# The population is validated in the same cell that fills it, because the two are one act: the
# expected set was written down before the first member ran, and `require_complete` is the only
# thing that turns it from a declaration into a published result. It can pass only when every
# planned model, checkpoint and portfolio size completed.
# %% tags=["results"]
# A sweep that recomputes everything and a sweep that recomputes nothing print the same summary
# unless the two are counted apart. `run_backtests` serves an identity that is already registered
# and complete instead of running it again, which is what makes a re-run affordable and what makes
# a bare member count say nothing about whether this run did any work.
#
# The runner already knows which it did and says so per member in `execution.diagnostics`, as
# `status` "reused" or "completed". Comparing against the registered hashes instead would be
# wrong in both directions: a registered-but-partial backtest is in that set, gets recomputed and
# would report as reused, and a preview re-run reads a table that excludes preview rows by default
# and would report every reused member as computed.
run_status: list[str] = []
backtests = []
for job in jobs:
execution = run_backtests(
study,
predictions=job["predictions"],
signal={"method": "equal_weight_top_k", "top_k": job["top_k"]},
chapter="16",
)
if len(execution.results) != job["expected"]:
raise RuntimeError("a baseline member disappeared during execution")
backtests.extend(execution.results)
run_status.extend(entry["status"] for entry in execution.diagnostics)
expected_count = sum(job["expected"] for job in jobs)
if len(backtests) != expected_count:
raise RuntimeError(f"expected {expected_count} baseline runs, found {len(backtests)}")
if {result.hash for result in backtests} != set(planned_hashes):
raise RuntimeError("completed baseline identities differ from the frozen plan")
if any(not result.complete for result in backtests):
raise RuntimeError("the baseline population contains an incomplete backtest")
backtest_rows = pl.DataFrame(
[
{
"backtest_hash": result.hash,
"prediction_hash": result.registry_record()["prediction_hash"],
"stage": result.registry_record()["stage"],
"complete": result.complete,
}
for result in backtests
]
)
if set(backtest_rows.get_column("stage")) != {"signal"}:
raise RuntimeError("equal-weight baseline runs must register with stage='signal'")
served = run_status.count("reused")
print(
f"Equal-weight baselines: {reuse_disclosure(len(backtests) - served, served)}, "
f"{len(backtests)} in the population"
)
if not include_preview:
if baseline_population is None:
raise RuntimeError("the canonical baseline population was not frozen before execution")
baseline_population.require_complete()
print(f"Official equal-weight population: {baseline_population.hash}")
backtest_rows
# %% [markdown]
# ## Key takeaways
#
# - The baseline evaluates every complete model configuration and checkpoint, so the population
# is configurations times checkpoints times portfolio sizes. That product is the number of
# trials every later comparison is drawn from.
# - Equal weight isolates the ranking's contribution. It is the reference the allocation, risk and
# cost stages are each measured against, which is why it is computed before any of them.
# - Catalog rows pass directly to the existing FX backtest engine.
# - Immutable population membership makes missing or failed jobs visible. Failures correlate with
# the thing being measured, so a sweep that published its survivors would flatter itself.
# - These rows carry `stage='signal'`. They are the baseline backtests; the stored value names the
# varied component, not a stage before the backtest.
#
# The Sharpe ratios here are validation numbers over a single history, and the best of a large
# grid is high partly because the grid is large. Nothing at this stage separates a configuration
# that ranks well from one that ranked well once. That separation is what the holdout is for, and
# it is spent once, at the end, against the whole candidate set rather than against the leader.
```출처의 라이선스에 따라 출처를 표시하고 전문을 공개합니다. 라이선스: MIT
이 요약은 원문을 바탕으로 Stratmill의 리서치 에이전트가 작성했으며, 원문을 복사한 것이 아닙니다.