跨案例特征筛选及其预测策略的局限
代码 《交易机器学习》
总结
本分析比较了九个交易案例研究中的特征筛选结果,并探讨通过筛选的特征能否预测表现更强的策略。筛选结合了 Benjamini–Hochberg 假发现控制,以及另一条基于案例研究效应量下限和各折一致性的路径。由于 PROCEED 可以通过任一路径得出,其计数不一定是 FDR 显著特征计数的子集。
本分析将各案例研究的 PROCEED 比例与所选配置的验证期和留出期夏普比率并列比较。文档提醒,观测数量较少,且已登记的回测并不完整,因此无法据此确定特征通过筛选是否能预测策略表现。文档还解释,单变量信息系数与联合模型预测回答的是不同问题:没有单个特征通过筛选,并不能证明市场不可预测。结果取决于重新生成的台账、各案例的特定阈值、相关特征检验,以及每项研究只选取一个配置。
核心观点
- PROCEED结合假发现显著性,以及基于效应量和各折稳定性的另一条筛选路径。
- FDR 份额采用统一阈值,而 PROCEED 比例反映各案例特定的下限;若不考虑这一点,两者无法直接比较。
- 单个特征的单变量信息系数无法衡量联合模型所用特征组合的价值。
- 有留出期回测的研究数量较少,无法确定特征通过筛选与策略夏普比率之间存在关系。
- 候选特征之间的相关性以及配置选择,限制了比较结果能够支持的结论。
标签
全文
# 02_feature_evaluation.py
```py
# ---
# jupyter:
# jupytext:
# cell_metadata_filter: tags,-all
# text_representation:
# extension: .py
# format_name: percent
# format_version: '1.3'
# jupytext_version: 1.19.3
# kernelspec:
# display_name: Python 3 (ipykernel)
# language: python
# name: python3
# ---
# %% [markdown]
# # Feature Evaluation Across Case Studies
#
# **Docker image**: `ml4t`
#
# Each case study runs an FDR-controlled triage pass on its candidate
# feature menu (Ch7 §7.3 sets the triage gates, §7.4 the FDR control),
# classifying every feature as PROCEED, REVISE,
# or STOP. This notebook reads those nine ledgers side by side and asks
# whether feature-level survival predicts strategy-level survival once
# the rest of the pipeline runs.
#
# The forward link is the point. §03 (signal quality) and §04
# (signal-to-strategy) build on whatever this notebook surfaces.
#
# **Book Reference**: Chapter 20, Section 20.2.
#
# **Prerequisites**: Run [`01_aggregate_synthesis`](01_aggregate_synthesis.ipynb)
# first. Each case study's `case_studies/{cs}/evaluation/triage_ledger.parquet`
# must exist (regenerated by the per-case-study `05_evaluation` notebooks).
# %%
"""Ch20 Feature Evaluation — cross-case-study triage ledger comparison."""
import matplotlib.pyplot as plt
import numpy as np
import polars as pl
from IPython.display import Markdown, display
from case_studies.utils.analytics import CASE_STUDY_IDS, SHORT_NAMES, load_triage_ledger
from case_studies.utils.strategy_analysis import rank_one
from utils.paths import REPO_ROOT, get_chapter_dir
from utils.style import COLORS, show_with_alt
# %% tags=["parameters"]
MAX_CASE_STUDIES = 0 # 0 = all available
# Benjamini-Hochberg level the ledger's `fdr_p` is compared against. The ledgers
# store the adjusted p-value rather than a significance flag, so the threshold
# lives here and is applied once.
FDR_ALPHA = 0.05
# %%
OUTPUT_DIR = get_chapter_dir(20) / "output"
CS_LIST = CASE_STUDY_IDS[:MAX_CASE_STUDIES] if MAX_CASE_STUDIES else CASE_STUDY_IDS
# %% [markdown]
# ## 1. Load Triage Ledgers
#
# Each per-case-study ledger has the same schema: one row per candidate feature,
# with IC moment estimates, a HAC-adjusted t-statistic, a Benjamini-Hochberg
# adjusted p-value (`fdr_p`), fold sign consistency, monotonicity, coverage, a
# categorical decision in {PROCEED, REVISE, STOP}, and a `note` giving the reason
# for that decision.
#
# The ledgers are generated artifacts, not repository content. Without them this
# notebook has nothing to compare, so a missing ledger stops the run rather than
# producing an empty table under a heading that promises nine case studies.
# %%
ledgers = {cs: load_triage_ledger(cs) for cs in CS_LIST}
available = [cs for cs, df in ledgers.items() if df is not None]
missing = [cs for cs in CS_LIST if cs not in available]
print(f"Loaded {len(available)}/{len(CS_LIST)} ledgers")
if missing:
raise FileNotFoundError(
f"No triage ledger for {missing}. Each is written by that case study's "
"`05_evaluation` notebook to case_studies/{cs}/evaluation/"
"triage_ledger.parquet; run those first."
)
# `fdr_sig` is derived here rather than read: the ledger stores the adjusted
# p-value, and a stored boolean would freeze one alpha into the artifact.
panel = pl.concat(
[
df.select(
"feature",
"decision",
"note",
"case_study",
fdr_sig=pl.col("fdr_p") < FDR_ALPHA,
)
for df in ledgers.values()
if df is not None
],
how="vertical_relaxed",
)
# %% [markdown]
# ## 2. Two Routes to PROCEED
#
# A feature earns PROCEED by either of two independent routes, recorded in the
# ledger's `note`: it clears Benjamini-Hochberg at the level above
# (`fdr_significant`), or its IC clears that case study's own effect-size floor
# and holds its sign across enough folds (`stable_and_above_threshold`) without
# clearing BH. The floor is set per case study, from 0.003 to 0.01 in absolute
# mean IC, so the second route is calibrated to each market rather than shared
# across them. The second
# route exists because BH over a menu of dozens of correlated features is a
# blunt instrument at these sample sizes, and a feature can be worth carrying
# forward on effect size and stability alone.
#
# PROCEED is therefore a union of the two routes, not a narrowing of the first.
# A case study can show more PROCEED features than FDR-significant ones, so the
# three counts below are drawn side by side rather than stacked -
# stacking them would assert a nesting that does not hold.
# %%
funnel = (
panel.group_by("case_study")
.agg(
n_candidate=pl.len(),
n_fdr_sig=pl.col("fdr_sig").sum(),
n_proceed=(pl.col("decision") == "PROCEED").sum(),
n_proceed_by_fdr=((pl.col("decision") == "PROCEED") & pl.col("fdr_sig")).sum(),
)
.with_columns(
n_proceed_by_stability=pl.col("n_proceed") - pl.col("n_proceed_by_fdr"),
pct_fdr_sig=(pl.col("n_fdr_sig") / pl.col("n_candidate") * 100).round(1),
pct_proceed=(pl.col("n_proceed") / pl.col("n_candidate") * 100).round(1),
cs_short=pl.col("case_study").replace(SHORT_NAMES),
)
.sort("pct_proceed", descending=True)
)
funnel.select(
"cs_short",
"n_candidate",
"n_fdr_sig",
"n_proceed",
"n_proceed_by_fdr",
"n_proceed_by_stability",
"pct_fdr_sig",
"pct_proceed",
)
# %%
fig, ax = plt.subplots(figsize=(9, 5))
y = np.arange(funnel.height)
height = 0.26
series = [
("n_candidate", "candidates", COLORS["silver_muted"]),
("n_fdr_sig", f"clears BH-FDR at {FDR_ALPHA:.2f}", COLORS["blue_light"]),
("n_proceed", "PROCEED", COLORS["blue"]),
]
for offset, (col, label, color) in zip((-height, 0.0, height), series):
ax.barh(
y + offset,
funnel[col].to_numpy(),
height=height,
color=color,
label=label,
edgecolor=COLORS["neutral"],
linewidth=0.4,
)
ax.set_yticks(y)
ax.set_yticklabels(funnel["cs_short"].to_numpy())
ax.invert_yaxis()
ax.set_xlabel("Number of features")
ax.set_title("Candidate features, FDR-significant, and PROCEED, by case study")
ax.legend(loc="lower right", frameon=False)
show_with_alt(
fig,
"Grouped horizontal bars per case study giving the candidate feature count, "
"the number clearing Benjamini-Hochberg, and the number reaching PROCEED. "
"The PROCEED bar exceeds the FDR-significant bar in most case studies, "
"because the stability route to PROCEED does not require BH significance.",
)
# %% [markdown]
# The nine case studies share the triage structure, and the two columns are
# comparable to different degrees. The BH-FDR share is recomputed here from each
# ledger's stored p-values at one `FDR_ALPHA`, so it is on a common scale. The
# PROCEED share is read from each ledger's own decision, taken at that case
# study's own effect-size floor and sign-consistency minimum, so it is not. The
# summary below is computed from the table above rather than typed, so it cannot
# drift from the ledgers as they are regenerated.
# %% tags=["results"]
_f = funnel.sort("pct_fdr_sig", descending=True)
_none = _f.filter(pl.col("n_fdr_sig") == 0)["cs_short"].to_list()
_top = _f.row(0, named=True)
_pr = funnel.sort("pct_proceed", descending=True)
_hi, _lo = _pr.row(0, named=True), _pr.row(-1, named=True)
_stab = funnel["n_proceed_by_stability"].sum()
_total_proceed = funnel["n_proceed"].sum()
display(
Markdown(
f"Across {funnel.height} case studies, the share of candidate features "
f"clearing BH-FDR at {FDR_ALPHA:.2f} runs from "
f"{_f['pct_fdr_sig'].min():.1f} to {_top['pct_fdr_sig']:.1f} percent "
f"({_top['cs_short']} highest). "
+ (
f"{len(_none)} clear it for none of their candidates: {', '.join(_none)}. "
if _none
else ""
)
+ f"PROCEED rates run from {_lo['pct_proceed']:.1f} percent "
f"({_lo['cs_short']}, {_lo['n_proceed']} of {_lo['n_candidate']}) to "
f"{_hi['pct_proceed']:.1f} percent ({_hi['cs_short']}, "
f"{_hi['n_proceed']} of {_hi['n_candidate']}). "
f"Of {_total_proceed} PROCEED decisions in total, {_stab} "
"were reached on effect size and fold stability without clearing BH, "
"which is why the PROCEED bars are not contained inside the FDR bars."
)
)
# %% [markdown]
# A case study can reach zero surviving features and still support a trained
# model. Univariate IC asks whether one feature predicts on its own; the model
# is fitted on all of them jointly and can use combinations that no single
# column carries. The two measurements disagree by construction, so a zero here
# does not by itself establish that the market is unpredictable.
# %% [markdown]
# ## 3. Forward Link: Feature Survival vs Strategy Survival
#
# Pair each case study's PROCEED rate against the selected configuration's
# validation Sharpe and holdout Sharpe from §01. The triage ledger sits upstream
# of every model and backtest, so this is the most direct available test of
# whether a richer surviving feature menu produces a stronger strategy.
#
# It is a weak test. Sharpe figures exist only for case studies with registered
# backtests, and §01 reports that several registries are empty while their
# rebuild is in progress. The pairing below therefore rests on a handful of
# points, which is too few to support a claim about the relationship in either
# direction. It is shown because the absence of a relationship at this sample
# size is itself worth seeing, not because it settles the question.
# %%
mq = pl.read_parquet(OUTPUT_DIR / "measurement_quality.parquet")
forward = (
funnel.select(["case_study", "cs_short", "pct_proceed"])
.join(
mq.select(
pl.col("cs_id").alias("case_study"),
"rank1_val_sharpe",
"holdout_sharpe",
),
on="case_study",
how="left",
)
.sort("pct_proceed", descending=True)
)
forward
# %%
fig, axes = plt.subplots(1, 2, figsize=(11, 4.2), sharey=False)
for ax, ycol, ylabel in zip(
axes,
["rank1_val_sharpe", "holdout_sharpe"],
["Carrier validation Sharpe", "Carrier holdout Sharpe"],
):
pts = forward.filter(pl.col(ycol).is_not_null()).to_pandas()
ax.scatter(
pts["pct_proceed"],
pts[ycol],
color=COLORS["blue"],
s=40,
edgecolor=COLORS["neutral"],
linewidth=0.5,
)
for _, row in pts.iterrows():
ax.annotate(
row["cs_short"],
(row["pct_proceed"], row[ycol]),
fontsize=8,
xytext=(4, 3),
textcoords="offset points",
)
ax.axhline(0, color=COLORS["neutral"], linestyle="--", linewidth=0.7)
ax.set_xlabel("% candidate features in PROCEED")
ax.set_ylabel(ylabel)
fig.suptitle("Feature survival against strategy Sharpe, by case study")
show_with_alt(
fig,
"Two scatter panels plotting each case study's percentage of PROCEED "
"features against its validation Sharpe on the left and its holdout Sharpe "
"on the right, points labelled by case study. Only case studies with "
"registered backtests appear, and they are too few to show a trend.",
)
# %% tags=["results"]
_paired = forward.filter(pl.col("holdout_sharpe").is_not_null())
_absent = forward.filter(pl.col("holdout_sharpe").is_null())["cs_short"].to_list()
# rank_one, not a one-key sort: the case study this picks is named in the sentence below,
# so a tie on holdout Sharpe would publish whichever row the join happened to emit first.
_best = (
rank_one(_paired, by="holdout_sharpe", name="cs_short").row(0, named=True)
if _paired.height
else None
)
display(
Markdown(
f"{_paired.height} of {forward.height} case studies carry a holdout "
"Sharpe; the rest have no registered backtests"
+ (f" ({', '.join(_absent)})" if _absent else "")
+ ". "
+ (
f"Among those that do, the highest holdout Sharpe is "
f"{_best['holdout_sharpe']:+.2f} ({_best['cs_short']}, "
f"{_best['pct_proceed']:.1f} percent PROCEED). "
if _best
else ""
)
+ "With this many points, no ordering of PROCEED rate against holdout "
"Sharpe is distinguishable from chance, and none is claimed here."
)
)
# %% [markdown]
# What the count of surviving features cannot tell you is how large those
# features are relative to the frictions of the market they trade in, or how
# they are turned into positions. Those two questions are what §03 and §04
# measure, and they are where the differences between these case studies
# actually appear.
# %% [markdown]
# ## Takeaways
#
# - The nine case studies share the triage structure but not all of its
# parameters. The BH-FDR column is re-thresholded here at one alpha and can be
# compared across markets; the PROCEED column carries each case study's own
# effect-size floor, which spans a factor of three, so a difference in PROCEED
# rate is a difference in market and in floor together and cannot be assigned
# to either alone.
# - PROCEED is a union of two routes, BH significance and effect-size-plus-fold-
# stability. Reading it as a stricter version of BH significance inverts the
# relationship and inflates how selective the screen appears.
# - Univariate IC and joint prediction answer different questions. A case study
# where no single feature clears BH can still train a usable model, and that
# is a property of the two tests rather than evidence about the market.
# - Whether feature survival predicts strategy survival is not settled here. Too
# few case studies currently carry holdout backtests for the comparison to
# discriminate, and this notebook says so rather than reading a trend into
# five points.
#
# ## Known Limitations
#
# - The ledgers are regenerated by each case study's `05_evaluation` notebook and
# are not versioned with this repository. Numbers here move when those are
# re-run; nothing in this notebook is pinned to a particular generation.
# - BH-FDR treats the candidate menu as one family of tests, but the candidates
# are heavily correlated - overlapping return windows, related volatility
# measures - so the effective number of independent tests is smaller than the
# row count and the adjustment is conservative. This is part of why the
# stability route to PROCEED exists.
# - The forward link compares one selected configuration per case study, not a
# distribution over configurations, so it carries the selection uncertainty
# §01 documents in `measurement_quality.parquet`.
```在遵守原作品许可的前提下,附作者信息全文展示。 许可协议: MIT
此摘要由 Stratmill 研究智能体根据原文撰写,并非原文副本。