複数の取引ケースにおける特徴量スクリーニングと戦略成績との関係
ノートブック Machine Learning for Trading
サマリー
この分析では、9つの取引ケーススタディにまたがる特徴量トリアージ台帳を比較します。Benjamini–Hochberg法による偽発見率の有意性と、ケース固有の効果量下限および分割間の符号安定性を組み合わせた、PROCEED判断への別経路とを区別します。どちらかの経路を満たせばよいため、PROCEED特徴量の数はFDR閾値を通過する数を超える場合があります。続いて各ケースのPROCEED割合を検証およびホールドアウトのシャープレシオと対応させ、特徴量が選別後に残ることが戦略の存続を予測するかを調べます。
証拠には明確な限界があります。台帳は再生成され、候補特徴量は相関しており、登録済みバックテストがあるケースはわずかです。ケース間のPROCEED率も異なる効果量下限を使うため、完全には比較できません。単変量の情報係数スクリーニングとモデル全体の有用性は異なる問いに答えます。単一特徴量がスクリーニングを通過しないことは、組み合わせたモデルに価値がない証明にはなりません。このノートブックは、特徴量が残ることがシャープレシオを予測すると主張せず、利用可能な点だけでは関係と偶然を区別できないとしています。
主なアイデア
- PROCEEDは、FDRの有意性、または十分な効果量と分割間の安定性によって判断されます。
- PROCEEDの件数はFDR有意の件数の部分集合ではないため、別々に比較する必要があります。
- ケース固有の効果量下限により、市場間のPROCEED率の直接比較には限界があります。
- 単変量の特徴量スクリーニングとモデル全体の性能は異なる性質を評価します。
- 特徴量の残存と戦略のシャープレシオとの関係を推定するには、登録済みホールドアウト・バックテストが不足しています。
タグ
全文
# Feature Evaluation Across Case Studies
# Feature Evaluation Across Case Studies
**Docker image**: `ml4t`
Each case study runs an FDR-controlled triage pass on its candidate
feature menu (Ch7 §7.3 sets the triage gates, §7.4 the FDR control),
classifying every feature as PROCEED, REVISE,
or STOP. This notebook reads those nine ledgers side by side and asks
whether feature-level survival predicts strategy-level survival once
the rest of the pipeline runs.
The forward link is the point. §03 (signal quality) and §04
(signal-to-strategy) build on whatever this notebook surfaces.
**Book Reference**: Chapter 20, Section 20.2.
**Prerequisites**: Run [`01_aggregate_synthesis`](01_aggregate_synthesis.ipynb)
first. Each case study's `case_studies/{cs}/evaluation/triage_ledger.parquet`
must exist (regenerated by the per-case-study `05_evaluation` notebooks).
```python
"""Ch20 Feature Evaluation — cross-case-study triage ledger comparison."""
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from IPython.display import Markdown, display
from case_studies.utils.analytics import CASE_STUDY_IDS, SHORT_NAMES, load_triage_ledger
from case_studies.utils.strategy_analysis import rank_one
from utils.paths import REPO_ROOT, get_chapter_dir
from utils.style import COLORS, show_with_alt
```
```python
MAX_CASE_STUDIES = 0 # 0 = all available
# Benjamini-Hochberg level the ledger's `fdr_p` is compared against. The ledgers
# store the adjusted p-value rather than a significance flag, so the threshold
# lives here and is applied once.
FDR_ALPHA = 0.05
```
```python
OUTPUT_DIR = get_chapter_dir(20) / "output"
CS_LIST = CASE_STUDY_IDS[:MAX_CASE_STUDIES] if MAX_CASE_STUDIES else CASE_STUDY_IDS
```
## 1. Load Triage Ledgers
Each per-case-study ledger has the same schema: one row per candidate feature,
with IC moment estimates, a HAC-adjusted t-statistic, a Benjamini-Hochberg
adjusted p-value (`fdr_p`), fold sign consistency, monotonicity, coverage, a
categorical decision in {PROCEED, REVISE, STOP}, and a `note` giving the reason
for that decision.
The ledgers are generated artifacts, not repository content. Without them this
notebook has nothing to compare, so a missing ledger stops the run rather than
producing an empty table under a heading that promises nine case studies.
```python
ledgers = {cs: load_triage_ledger(cs) for cs in CS_LIST}
available = [cs for cs, df in ledgers.items() if df is not None]
missing = [cs for cs in CS_LIST if cs not in available]
print(f"Loaded {len(available)}/{len(CS_LIST)} ledgers")
if missing:
raise FileNotFoundError(
f"No triage ledger for {missing}. Each is written by that case study's "
"`05_evaluation` notebook to case_studies/{cs}/evaluation/"
"triage_ledger.parquet; run those first."
)
# `fdr_sig` is derived here rather than read: the ledger stores the adjusted
# p-value, and a stored boolean would freeze one alpha into the artifact.
panel = pl.concat(
[
df.select(
"feature",
"decision",
"note",
"case_study",
fdr_sig=pl.col("fdr_p") < FDR_ALPHA,
)
for df in ledgers.values()
if df is not None
],
how="vertical_relaxed",
)
```
## 2. Two Routes to PROCEED
A feature earns PROCEED by either of two independent routes, recorded in the
ledger's `note`: it clears Benjamini-Hochberg at the level above
(`fdr_significant`), or its IC clears that case study's own effect-size floor
and holds its sign across enough folds (`stable_and_above_threshold`) without
clearing BH. The floor is set per case study, from 0.003 to 0.01 in absolute
mean IC, so the second route is calibrated to each market rather than shared
across them. The second
route exists because BH over a menu of dozens of correlated features is a
blunt instrument at these sample sizes, and a feature can be worth carrying
forward on effect size and stability alone.
PROCEED is therefore a union of the two routes, not a narrowing of the first.
A case study can show more PROCEED features than FDR-significant ones, so the
three counts below are drawn side by side rather than stacked -
stacking them would assert a nesting that does not hold.
```python
funnel = (
panel.group_by("case_study")
.agg(
n_candidate=pl.len(),
n_fdr_sig=pl.col("fdr_sig").sum(),
n_proceed=(pl.col("decision") == "PROCEED").sum(),
n_proceed_by_fdr=((pl.col("decision") == "PROCEED") & pl.col("fdr_sig")).sum(),
)
.with_columns(
n_proceed_by_stability=pl.col("n_proceed") - pl.col("n_proceed_by_fdr"),
pct_fdr_sig=(pl.col("n_fdr_sig") / pl.col("n_candidate") * 100).round(1),
pct_proceed=(pl.col("n_proceed") / pl.col("n_candidate") * 100).round(1),
cs_short=pl.col("case_study").replace(SHORT_NAMES),
)
.sort("pct_proceed", descending=True)
)
funnel.select(
"cs_short",
"n_candidate",
"n_fdr_sig",
"n_proceed",
"n_proceed_by_fdr",
"n_proceed_by_stability",
"pct_fdr_sig",
"pct_proceed",
)
```
```python
fig, ax = plt.subplots(figsize=(9, 5))
y = np.arange(funnel.height)
height = 0.26
series = [
("n_candidate", "candidates", COLORS["silver_muted"]),
("n_fdr_sig", f"clears BH-FDR at {FDR_ALPHA:.2f}", COLORS["blue_light"]),
("n_proceed", "PROCEED", COLORS["blue"]),
]
for offset, (col, label, color) in zip((-height, 0.0, height), series):
ax.barh(
y + offset,
funnel[col].to_numpy(),
height=height,
color=color,
label=label,
edgecolor=COLORS["neutral"],
linewidth=0.4,
)
ax.set_yticks(y)
ax.set_yticklabels(funnel["cs_short"].to_numpy())
ax.invert_yaxis()
ax.set_xlabel("Number of features")
ax.set_title("Candidate features, FDR-significant, and PROCEED, by case study")
ax.legend(loc="lower right", frameon=False)
show_with_alt(
fig,
"Grouped horizontal bars per case study giving the candidate feature count, "
"the number clearing Benjamini-Hochberg, and the number reaching PROCEED. "
"The PROCEED bar exceeds the FDR-significant bar in most case studies, "
"because the stability route to PROCEED does not require BH significance.",
)
```
The nine case studies share the triage structure, and the two columns are
comparable to different degrees. The BH-FDR share is recomputed here from each
ledger's stored p-values at one `FDR_ALPHA`, so it is on a common scale. The
PROCEED share is read from each ledger's own decision, taken at that case
study's own effect-size floor and sign-consistency minimum, so it is not. The
summary below is computed from the table above rather than typed, so it cannot
drift from the ledgers as they are regenerated.
```python
_f = funnel.sort("pct_fdr_sig", descending=True)
_none = _f.filter(pl.col("n_fdr_sig") == 0)["cs_short"].to_list()
_top = _f.row(0, named=True)
_pr = funnel.sort("pct_proceed", descending=True)
_hi, _lo = _pr.row(0, named=True), _pr.row(-1, named=True)
_stab = funnel["n_proceed_by_stability"].sum()
_total_proceed = funnel["n_proceed"].sum()
display(
Markdown(
f"Across {funnel.height} case studies, the share of candidate features "
f"clearing BH-FDR at {FDR_ALPHA:.2f} runs from "
f"{_f['pct_fdr_sig'].min():.1f} to {_top['pct_fdr_sig']:.1f} percent "
f"({_top['cs_short']} highest). "
+ (
f"{len(_none)} clear it for none of their candidates: {', '.join(_none)}. "
if _none
else ""
)
+ f"PROCEED rates run from {_lo['pct_proceed']:.1f} percent "
f"({_lo['cs_short']}, {_lo['n_proceed']} of {_lo['n_candidate']}) to "
f"{_hi['pct_proceed']:.1f} percent ({_hi['cs_short']}, "
f"{_hi['n_proceed']} of {_hi['n_candidate']}). "
f"Of {_total_proceed} PROCEED decisions in total, {_stab} "
"were reached on effect size and fold stability without clearing BH, "
"which is why the PROCEED bars are not contained inside the FDR bars."
)
)
```
A case study can reach zero surviving features and still support a trained
model. Univariate IC asks whether one feature predicts on its own; the model
is fitted on all of them jointly and can use combinations that no single
column carries. The two measurements disagree by construction, so a zero here
does not by itself establish that the market is unpredictable.
## 3. Forward Link: Feature Survival vs Strategy Survival
Pair each case study's PROCEED rate against the selected configuration's
validation Sharpe and holdout Sharpe from §01. The triage ledger sits upstream
of every model and backtest, so this is the most direct available test of
whether a richer surviving feature menu produces a stronger strategy.
It is a weak test. Sharpe figures exist only for case studies with registered
backtests, and §01 reports that several registries are empty while their
rebuild is in progress. The pairing below therefore rests on a handful of
points, which is too few to support a claim about the relationship in either
direction. It is shown because the absence of a relationship at this sample
size is itself worth seeing, not because it settles the question.
```python
mq = pl.read_parquet(OUTPUT_DIR / "measurement_quality.parquet")
forward = (
funnel.select(["case_study", "cs_short", "pct_proceed"])
.join(
mq.select(
pl.col("cs_id").alias("case_study"),
"rank1_val_sharpe",
"holdout_sharpe",
),
on="case_study",
how="left",
)
.sort("pct_proceed", descending=True)
)
forward
```
```python
fig, axes = plt.subplots(1, 2, figsize=(11, 4.2), sharey=False)
for ax, ycol, ylabel in zip(
axes,
["rank1_val_sharpe", "holdout_sharpe"],
["Carrier validation Sharpe", "Carrier holdout Sharpe"],
):
pts = forward.filter(pl.col(ycol).is_not_null()).to_pandas()
ax.scatter(
pts["pct_proceed"],
pts[ycol],
color=COLORS["blue"],
s=40,
edgecolor=COLORS["neutral"],
linewidth=0.5,
)
for _, row in pts.iterrows():
ax.annotate(
row["cs_short"],
(row["pct_proceed"], row[ycol]),
fontsize=8,
xytext=(4, 3),
textcoords="offset points",
)
ax.axhline(0, color=COLORS["neutral"], linestyle="--", linewidth=0.7)
ax.set_xlabel("% candidate features in PROCEED")
ax.set_ylabel(ylabel)
fig.suptitle("Feature survival against strategy Sharpe, by case study")
show_with_alt(
fig,
"Two scatter panels plotting each case study's percentage of PROCEED "
"features against its validation Sharpe on the left and its holdout Sharpe "
"on the right, points labelled by case study. Only case studies with "
"registered backtests appear, and they are too few to show a trend.",
)
```
```python
_paired = forward.filter(pl.col("holdout_sharpe").is_not_null())
_absent = forward.filter(pl.col("holdout_sharpe").is_null())["cs_short"].to_list()
# rank_one, not a one-key sort: the case study this picks is named in the sentence below,
# so a tie on holdout Sharpe would publish whichever row the join happened to emit first.
_best = (
rank_one(_paired, by="holdout_sharpe", name="cs_short").row(0, named=True)
if _paired.height
else None
)
display(
Markdown(
f"{_paired.height} of {forward.height} case studies carry a holdout "
"Sharpe; the rest have no registered backtests"
+ (f" ({', '.join(_absent)})" if _absent else "")
+ ". "
+ (
f"Among those that do, the highest holdout Sharpe is "
f"{_best['holdout_sharpe']:+.2f} ({_best['cs_short']}, "
f"{_best['pct_proceed']:.1f} percent PROCEED). "
if _best
else ""
)
+ "With this many points, no ordering of PROCEED rate against holdout "
"Sharpe is distinguishable from chance, and none is claimed here."
)
)
```
What the count of surviving features cannot tell you is how large those
features are relative to the frictions of the market they trade in, or how
they are turned into positions. Those two questions are what §03 and §04
measure, and they are where the differences between these case studies
actually appear.
## Takeaways
- The nine case studies share the triage structure but not all of its
parameters. The BH-FDR column is re-thresholded here at one alpha and can be
compared across markets; the PROCEED column carries each case study's own
effect-size floor, which spans a factor of three, so a difference in PROCEED
rate is a difference in market and in floor together and cannot be assigned
to either alone.
- PROCEED is a union of two routes, BH significance and effect-size-plus-fold-
stability. Reading it as a stricter version of BH significance inverts the
relationship and inflates how selective the screen appears.
- Univariate IC and joint prediction answer different questions. A case study
where no single feature clears BH can still train a usable model, and that
is a property of the two tests rather than evidence about the market.
- Whether feature survival predicts strategy survival is not settled here. Too
few case studies currently carry holdout backtests for the comparison to
discriminate, and this notebook says so rather than reading a trend into
five points.
## Known Limitations
- The ledgers are regenerated by each case study's `05_evaluation` notebook and
are not versioned with this repository. Numbers here move when those are
re-run; nothing in this notebook is pinned to a particular generation.
- BH-FDR treats the candidate menu as one family of tests, but the candidates
are heavily correlated - overlapping return windows, related volatility
measures - so the effective number of independent tests is smaller than the
row count and the adjustment is conservative. This is part of why the
stability route to PROCEED exists.
- The forward link compares one selected configuration per case study, not a
distribution over configurations, so it carries the selection uncertainty
§01 documents in `measurement_quality.parquet`.

出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。