ETFアンサンブルのICが高くても、ポートフォリオが劣る理由
コード Machine Learning for Trading
サマリー
この実験では、勾配ブースティング木、表形式ニューラルネットワーク、潜在因子モデルによる月次ETF順位予測の検証データ予測を組み合わせます。まず予測をクロスセクションでzスコア標準化し、その後に平均します。異分散・自己相関を考慮した不確実性推定付きの情報係数(IC)と、少数のETFを選んで加重するコスト考慮型のロングオンリー・バックテストでアンサンブルを評価します。
アンサンブルは集計ICがLSTMベンチマークより高い一方、シャープレシオは低く、最大ドローダウンは大きくなっています。分割ごとの結果を見るとその理由が分かります。アンサンブルの順位相関は大きく変動し、負の分割もある一方、LSTMの順位相関は一貫して正です。この資料は、平均的な順位付け精度が高くても、ポートフォリオのリターンが良くなるとは限らないことを示します。特に予測の不安定さがスコアに基づく配分に影響する場合はそうです。比較は特定の検証期間と戦略設定に限られ、この結果からアンサンブルが一般に成績を悪化させるとは言えません。
主なアイデア
- クロスセクションのzスコアを平均することで、尺度の異なる予測を組み合わせられます。
- 報告されたバックテストでは、アンサンブルはLSTMベンチマークより集計ICを改善しましたが、シャープレシオは低くなっています。
- 分割ごとのICが負になったり大きく変動したりすると、平均的な順位精度が高くてもポートフォリオ成績が損なわれる場合があります。
- 予測の質と実現した戦略成績との関係は、ポートフォリオ構築によって変わります。
タグ
全文
# run_etfs_20260504T223150.json
```json
{
"model": "deepseek/deepseek-v4-pro",
"case_study": "etfs",
"sandbox": "/tmp/etfs_operator_sandbox_20260504T221051",
"iterations": 39,
"total_in_tokens": 2441130,
"total_out_tokens": 19492,
"final_summary": "## \u00a720.9 Next Step: Ensemble GBM + Tabular DL + CAE \u2014 Result\n\n### Experiment\nEnsembled the rank-1 validation prediction sets for GBM (`leaves_7_mse`), tabular DL (`tabm_l`), and latent factors (`cae`) on label `fwd_ret_21d`. Combined via z-score cross-sectional averaging (preserves rank ordering across models with different prediction scales). Evaluated IC with HAC standard errors via `ml4t-diagnostic` (`cross_sectional_ic_series` + `compute_ic_hac_stats`), then ran cost-aware backtests matching the LSTM baseline's `score_weighted_top_k` / `top_k=20` / monthly cadence / long-only configuration.\n\n### Skills Used\n- `information-coefficient` \u2014 IC with HAC SE, IC_IR, pct_positive\n- `evaluate-factor` \u2014 per-fold IC stability analysis\n- `deflated-sharpe` \u2014 PSR for backtest significance\n- `backtest-overfitting` \u2014 PBO awareness\n- `registry-system` \u2014 content-addressed prediction discovery\n\n### Results (LSTM date range: 2016-03-24 \u2192 2023-12-29, score_weighted_top_k / top_k=20)\n\n| Model | IC | t(HAC) | Sharpe | Sharpe CI95 | PSR p | MaxDD |\n|-------|-----|--------|--------|-------------|-------|-------|\n| LSTM_h64 (baseline) | 0.0521 | 3.40 | **0.922** | [0.37, 1.53] | 0.005 | -0.179 |\n| Ensemble | 0.0649 | 4.60 | **0.562** | [-0.03, 1.25] | 0.042 | -0.367 |\n\n**Sharpe delta vs LSTM baseline: -0.36** (ensemble underperforms).\n\n### Diagnosis\nThe ensemble achieves the highest IC (0.0649, t=4.60) of any model \u2014 better than LSTM's 0.0521. However, per-fold IC reveals high variance: Folds 0, 5, 7 are negative (IC \u2248 -0.03 to -0.08), while Folds 3-4 are very strong (IC \u2248 0.16-0.19). The LSTM has all-positive folds (0.004\u20130.106) with lower magnitude but greater consistency. This fold instability in the ensemble translates to worse portfolio construction: the `score_weighted_top_k` scheme amplifies signal differences, and negative-fold predictions drag down the strategy.\n\n### Recommendation\n**No improvement.** The ensemble does not stabilize holdout Sharpe. The LSTM remains the rank-1 configuration. The IC-vs-Sharpe gap (better rank correlation, worse portfolio performance) is the canonical teaching point from the README: \"the family with the highest rank correlation is not the family with the highest portfolio Sharpe.\" The ensemble's cross-model averaging improves IC but introduces fold-level instability that portfolio construction cannot absorb.",
"trace": [
{
"turn": 1,
"type": "tool_call",
"name": "read_file",
"args": {
"path": "/tmp/etfs_operator_sandbox_20260504T221051/etfs/README.md"
},
"result": {
"path": "/tmp/etfs_operator_sandbox_20260504T221051/etfs/README.md",
"size_bytes": 6476,
"truncated": false,
"content": "# Case Study: ETF Cross-Asset Exposures\n\nThis case study applies the ML4T workflow to 100 exchange-traded funds spanning equities, fixed income, commodities, currencies, and real estate. ETFs offer a clean laboratory for cross-asset rotation: standardized pricing, deep liquidity, and broad asset-class coverage at a single rebalance cadence.\n\nThe configuration is the most cost-favorable in the book \u2014 long-only rank-and-rebalance, monthly month-end decisions on a 21-day forward-return label, with a 5--15 bps-per-leg cost model. That cadence makes it the natural setting for the broadest model-family comparison in the book: linear, GBM, tabular DL, sequence DL, latent factors, and causal DML are all trained on the same feature panel. The teaching point is the gap between IC and Sharpe \u2014 the family with the highest rank correlation is not the family with the highest portfolio Sharpe \u2014 which makes ETFs the canonical setting for the \"portfolio construction mediates prediction quality\" thread that runs through Ch16--Ch20.\n\n## At a Glance\n\n| Property | Value |\n|----------|-------|\n| Asset Class | Multi-asset ETFs |\n| Frequency | Daily data, monthly decisions |\n| Universe | 100 ETFs across 9 categories |\n| History | 2006--2025 |\n| Primary Label | fwd_ret_21d |\n| CV Folds | 8 (10Y train, 1Y val) |\n| Cost Model | Material (5--15 bps per leg) |\n\n## Pipeline\n\n| Stage | Notebook | Chapter | Description |\n|-------|----------|---------|-------------|\n| Setup | [`01_setup`](01_setup.ipynb) | Ch6 | Universe definition, monthly rebalance cadence, long-only cost model |\n| Labels | [`02_labels`](02_labels.ipynb) | Ch7 | 21-day and 5-day forward returns with walk-forward splits |\n| Features | [`03_financial_features`](03_financial_features.ipynb) | Ch8 | Momentum, volatility, and cross-asset ranking features |\n| Temporal | [`04_model_based_features`](04_model_based_features.ipynb) | Ch9 | ARIMA, HMM, and spectral features from walk-forward fits |\n| Evaluation | [`05_evaluation`](05_evaluation.ipynb) | Ch7--9 | Feature-label IC diagnostics across all engineered features |\n| Linear | [`06_linear`](06_linear.ipynb) | Ch11 | Ridge, LASSO, ElasticNet baseline for cross-asset momentum |\n| GBM | [`07_gbm`](07_gbm.ipynb) | Ch12 | LightGBM with Optuna testing non-linear interactions |\n| Tabular DL | [`08_tabular_dl`](08_tabular_dl.ipynb) | Ch12 | TabM rank-1 adapter MLP ensemble |\n| LSTM | [`09_dl_lstm`](09_dl_lstm.ipynb) | Ch13 | Temporal gating over sequential ETF return windows |\n| TSMixer | [`10_dl_tsmixer`](10_dl_tsmixer.ipynb) | Ch13 | Cross-asset lead-lag patterns via time-feature mixing |\n| Latent Factors | [`11_latent_factors`](11_latent_factors.ipynb) | Ch14 | Factor extraction across the ETF universe |\n| Causal DML | [`12_causal_dml`](12_causal_dml.ipynb) | Ch15 | Does momentum cause future ETF returns or reflect confounders? |\n| Model Analysis | [`13_model_analysis`](13_model_analysis.ipynb) | Ch11--15 | Cross-family IC comparison, checkpoint sensitivity, fold stability |\n| Backtest | [`14_backtest`](14_backtest.ipynb) | Ch16 | Strategy simulation with falsification against equal-weight |\n| Portfolio | [`15_portfolio_management`](15_portfolio_management.ipynb) | Ch17 | Score-weighted, risk-parity, and mean-variance allocation |\n| Costs | [`16_costs`](16_costs.ipynb) | Ch18 | Transaction cost impact on the momentum edge |\n| Risk | [`17_risk_management`](17_risk_management.ipynb) | Ch19 | Drawdown controls and regime-conditional position sizing |\n| Strategy Analysis | [`18_strategy_analysis`](18_strategy_analysis.ipynb) | Ch20 | End-to-end strategy assessment with IC, Sharpe, and cost analysis |\n\n## Key Results\n\n**Signal quality**: Daily-pooled IC for the highest-Sharpe configuration is +0.052 [+0.009, +0.095] (HAC $t=2.37$, $p=0.018$, excludes zero on the positive side); pct-positive is 56.4%, modestly above coin flip. The rank-correlation prior is statistically resolved at the validation window, with magnitude small but credibly nonzero.\n\n**Strategy-stage performance with CIs**: The signal-stage configuration with the highest validation Sharpe is `deep_learning/lstm_h64` on `fwd_ret_21d`. Validation Sharpe is +0.92 [+0.40, +1.49], PSR $p=0.005$ \u2014 both Sharpe CI and PSR exclude zero on the positive side. PBO is 0.443 across 8 folds \u00d7 70 combinations, flagging meaningful overfit probability without a deflated-Sharpe value to anchor magnitude on this row.\n\n**Holdout closure**: Validation\u2192holdout Sharpe difference is $-0.53$ [$-1.92$, $+1.03$] ($p=0.475$, straddles zero \u2014 the 502-day holdout is too short relative to bootstrap spread to resolve decay vs zero). Against an equal-weight benchmark, the holdout-period Sharpe difference is $-0.68$ [$-1.68$, $+0.12$] ($p=0.124$); the two-sided test does not reject, but `prob_challenger_wins` is 0.043, indicating strong directional underperformance driven by the 2024--2025 broad-equity rally where cross-asset rotation toward bonds and commodities gave back ground to a static equity-weighted universe.\n\n**Friction floor**: Cost sensitivity scans 11 levels from 0 to 50 bps per leg. The highest-Sharpe configuration stays positive across the full grid; median Sharpe across all configurations stays positive through realistic ETF friction (\u22645 bps). Both kill gates pass \u2014 validation Sharpe lower bound \u2265 0, and holdout strategy CI does not exclude zero negatively.\n\n## Running\n\n```bash\n# From repo root\nuv run python case_studies/etfs/01_setup.py\nuv run python case_studies/etfs/02_labels.py\nuv run python case_studies/etfs/03_financial_features.py\nuv run python case_studies/etfs/04_model_based_features.py\nuv run python case_studies/etfs/05_evaluation.py\nuv run python case_studies/etfs/06_linear.py\nuv run python case_studies/etfs/07_gbm.py\nuv run python case_studies/etfs/08_tabular_dl.py\nuv run python case_studies/etfs/09_dl_lstm.py\nuv run python case_studies/etfs/10_dl_tsmixer.py\nuv run python case_studies/etfs/11_latent_factors.py\nuv run python case_studies/etfs/12_causal_dml.py\nuv run python case_studies/etfs/13_model_analysis.py\nuv run python case_studies/etfs/14_backtest.py\nuv run python case_studies/etfs/15_portfolio_management.py\nuv run python case_studies/etfs/16_costs.py\nuv run python case_studies/etfs/17_risk_management.py\nuv run python case_studies/etfs/18_strategy_analysis.py\n```\n\n## Run Log\n\nModel training runs, predictions, and backtest results are tracked in a content-addressed registry under `run_log/registry.db`.\n"
}
},
{
"turn": 1,
"type": "tool_call",
"name": "list_skills",
"args": {
"category": "validation"
},
"result": {
"category": "validation",
"n_skills": 8,
"skills": [
{
"name": "ml4t-cpcv",
"category": "validation",
"path": "~/ml4t/skills/validation/cpcv/SKILL.md",
"description": "Combinatorial Purged CV generates a distribution of backtest paths instead of a single estimate. Use when quantifying strategy robustness and overfitting probability.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-deflated-sharpe",
"category": "validation",
"path": "~/ml4t/skills/validation/deflated-sharpe/SKILL.md",
"description": "Adjust the Sharpe ratio for multiple testing bias when selecting from many trials. Use when reporting strategy performance after parameter or model search.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-drift-detection",
"category": "validation",
"path": "~/ml4t/skills/validation/drift-detection/SKILL.md",
"description": "Detect when model inputs or predictions shift from the training distribution using PSI and statistical tests. Use when monitoring deployed models or diagnosing performance degradation.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-evaluate-factor",
"category": "validation",
"path": "~/ml4t/skills/validation/evaluate-factor/SKILL.md",
"description": "Evaluate alpha factor quality with IC analysis, quantile spreads, turnover, and decay. Use when deciding whether a signal has enough predictive power to trade.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-purging-embargo",
"category": "validation",
"path": "~/ml4t/skills/validation/purging-embargo/SKILL.md",
"description": "Remove training samples whose labels overlap the test period and add an embargo buffer. Use when performing time-series CV to prevent leakage between folds.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-shap-analysis",
"category": "validation",
"path": "~/ml4t/skills/validation/shap-analysis/SKILL.md",
"description": "Explain model predictions with SHAP values instead of biased built-in feature importance. Use when interpreting which features drive model decisions.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-stationarity-tests",
"category": "validation",
"path": "~/ml4t/skills/validation/stationarity-tests/SKILL.md",
"description": "Test whether time-series features are stationary using ADF and KPSS tests. Use when selecting or transforming features to verify stationarity assumptions for predictive models.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-walk-forward-cv",
"category": "validation",
"path": "~/ml4t/skills/validation/walk-forward-cv/SKILL.md",
"description": "Rolling or expanding window CV that preserves temporal order. Use when evaluating ML models on time-series data where standard k-fold causes temporal leakage.",
"library": "ml4t-diagnostic"
}
]
}
},
{
"turn": 2,
"type": "tool_call",
"name": "list_skills",
"args": {
"category": "concepts"
},
"result": {
"category": "concepts",
"n_skills": 10,
"skills": [
{
"name": "ml4t-backtest-overfitting",
"category": "concepts",
"path": "~/ml4t/skills/concepts/backtest-overfitting/SKILL.md",
"description": "Detect and prevent overfitting to historical data via multiple testing corrections and pre-registration. Use when evaluating strategy variants to ensure performance is not a data-mining artifact.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-causal-identification",
"category": "concepts",
"path": "~/ml4t/skills/concepts/causal-identification/SKILL.md",
"description": "Validate causal claims using DAG adjustment sets, bad-control detection, and refutation tests. Use when distinguishing genuine factor effects from confounded associations.",
"library": ""
},
{
"name": "ml4t-data-leakage",
"category": "concepts",
"path": "~/ml4t/skills/concepts/data-leakage/SKILL.md",
"description": "Prevent train-test contamination, target leakage, and temporal leakage. Use when splitting data, fitting preprocessors, or engineering features for time-series ML.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-information-coefficient",
"category": "concepts",
"path": "~/ml4t/skills/concepts/information-coefficient/SKILL.md",
"description": "Measure predictive signal quality with IC, Rank IC, and IC_IR. Use when evaluating whether a feature has predictive power for returns.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-lookahead-bias",
"category": "concepts",
"path": "~/ml4t/skills/concepts/lookahead-bias/SKILL.md",
"description": "Detect future information leaking into features, labels, or evaluation. Use when any pipeline step might expose data not yet available at prediction time.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-non-stationarity",
"category": "concepts",
"path": "~/ml4t/skills/concepts/non-stationarity/SKILL.md",
"description": "Handle changing statistical properties in financial time series. Use when features or model performance degrade over time.",
"library": "ml4t-diagnostic"
},
{
"name": "ml4t-point-in-time",
"category": "concepts",
"path": "~/ml4t/skills/concepts/point-in-time/SKILL.md",
"description": "Ensure data reflects what was known at each decision point, not revised or restated values. Use when joining fundamental, macro, or alternative data to price series.",
"library": "ml4t-data"
},
{
"name": "ml4t-regime-awareness",
"category": "concepts",
"path": "~/ml4t/skills/concepts/regime-awareness/SKILL.md",
"description": "Market regimes as conditioning features for risk scaling, not timing signals. Use when incorporating regime detection into strategy logic.",
"library": "ml4t-engineer"
},
{
"name": "ml4t-survivorship-bias",
"category": "concepts",
"path": "~/ml4t/skills/concepts/survivorship-bias/SKILL.md",
"description": "Account for delisted and removed securities in historical analysis. Use when constructing universes or computing cross-sectional features to avoid survivor-only inflation.",
"library": "ml4t-data"
},
{
"name": "ml4t-transaction-costs",
"category": "concepts",
"path": "~/ml4t/skills/concepts/transaction-costs/SKILL.md",
"description": "Estimate whether a strategy can survive spread, slippage, and market impact before full simulation. Use when screening strategy feasibility early.",
"library": ""
}
]
}
},
{
"turn": 2,
"type": "tool_call",
"name": "list_skills",
"args": {
"category": "infrastructure"
},
"result": {
"category": "infrastructure",
"n_skills": 4,
"skills": [
{
"name": "ml4t-canonical-schema",
"category": "infrastructure",
"path": "~/ml4t/skills/infrastructure/canonical-schema/SKILL.md",
"description": "Standardized data schema across all financial datasets. Use when defining or enforcing column names, types, and index conventions.",
"library": "ml4t-data"
},
{
"name": "ml4t-case-study-pipeline",
"category": "infrastructure",
"path": "~/ml4t/skills/infrastructure/case-study-pipeline/SKILL.md",
"description": "Filesystem and artifact-contract pattern for reproducible case studies. Use when organizing a research project for reproducibility and collaboration.",
"library": ""
},
{
"name": "ml4t-polars-patterns",
"category": "infrastructure",
"path": "~/ml4t/skills/infrastructure/polars-patterns/SKILL.md",
"description": "Polars-first data processing patterns for financial data. Use when writing efficient grouped, windowed, or lazy-evaluated data transformations.",
"library": ""
},
{
"name": "ml4t-registry-system",
"category": "infrastructure",
"path": "~/ml4t/skills/infrastructure/registry-system/SKILL.md",
"description": "Content-addressed experiment tracking for ML trading models. Use when versioning models, features, or experiment artifacts for reproducibility.",
"library": ""
}
]
}
},
{
"turn": 2,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT DISTINCT model_family, model_name, label, prediction_set_id, split_kind FROM prediction_sets ORDER BY model_family, model_name, split_kind"
},
"result": {
"error": "OperationalError: no such column: model_family"
}
},
{
"turn": 3,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT * FROM prediction_sets LIMIT 1"
},
"result": {
"columns": [
"prediction_hash",
"training_hash",
"checkpoint_value",
"checkpoint_kind",
"split",
"created_at"
],
"rows": [
{
"prediction_hash": "d533746293bd",
"training_hash": "1957228a8dcb",
"checkpoint_value": null,
"checkpoint_kind": null,
"split": "validation",
"created_at": "2026-04-10T11:16:46.462035+00:00"
}
],
"n_rows_returned": 1,
"truncated_at_200": false
}
},
{
"turn": 3,
"type": "tool_call",
"name": "read_skill",
"args": {
"name_or_path": "ml4t-information-coefficient"
},
"result": {
"error": "no skill matches 'ml4t-information-coefficient'",
"hint": "use list_skills to discover names"
}
},
{
"turn": 3,
"type": "tool_call",
"name": "read_skill",
"args": {
"name_or_path": "ml4t-registry-system"
},
"result": {
"error": "no skill matches 'ml4t-registry-system'",
"hint": "use list_skills to discover names"
}
},
{
"turn": 4,
"type": "tool_call",
"name": "read_skill",
"args": {
"name_or_path": "~/ml4t/skills/concepts/information-coefficient/SKILL.md"
},
"result": {
"path": "~/ml4t/skills/concepts/information-coefficient/SKILL.md",
"size_bytes": 4196,
"content": "---\nname: ml4t-information-coefficient\ndescription: \"Measure predictive signal quality with IC, Rank IC, and IC_IR. Use when evaluating whether a feature has predictive power for returns.\"\nwhen_to_use: \"Use when evaluating alpha signals, comparing features, or connecting signal strength to portfolio performance via the Fundamental Law\"\ndependencies: []\nmetadata:\n book_chapters: \"7, 8\"\n library: \"ml4t-diagnostic\"\n---\n# Information Coefficient\n\nIC is the correlation between a predicted signal and realized returns. It is the primary metric for judging whether a signal has predictive power before building a full backtest.\n\n## The Problem\n\nReporting a single IC value (or worse, the best IC from many trials) tells you almost nothing. IC varies over time and across market regimes. A signal with mean IC = 0.04 and std = 0.02 (IC_IR = 2.0) is far more valuable than one with mean IC = 0.08 and std = 0.10 (IC_IR = 0.8). Without time-series statistics and proper standard errors, you cannot distinguish a real signal from noise.\n\n## The Pattern\n\n### WRONG\n\n```python\nfrom scipy.stats import spearmanr\n\n# Single pooled IC -- hides time variation, inflates significance\nic, pval = spearmanr(all_predictions.flatten(), all_returns.flatten())\nprint(f\"IC = {ic:.4f}, p = {pval:.4f}\")\n```\n\n### CORRECT\n\n```python\nimport numpy as np\nfrom scipy.stats import spearmanr\nfrom statsmodels.stats.stattools import durbin_watson\n\n# Cross-sectional IC per period\nic_series = []\nfor t in timestamps:\n mask = dates == t\n if mask.sum() >= 10:\n ic, _ = spearmanr(predictions[mask], returns[mask])\n ic_series.append(ic)\n\nic_series = np.array(ic_series)\nic_mean = ic_series.mean()\nic_std = ic_series.std()\nic_ir = ic_mean / ic_std\n\n# Newey-West HAC standard error for significance\nfrom statsmodels.regression.linear_model import OLS\nfrom statsmodels.tools import add_constant\nols = OLS(ic_series, add_constant(np.ones(len(ic_series)))).fit(\n cov_type=\"HAC\", cov_kwds={\"maxlags\": 5}\n)\nt_stat = ols.tvalues[0]\n\nprint(f\"IC: {ic_mean:.4f}, IC_IR: {ic_ir:.2f}, t(HAC): {t_stat:.2f}\")\n```\n\n## IC Interpretation\n\n| IC range | Quality | Notes |\n|----------|---------|-------|\n| > 0.10 | Excellent | Rare; verify no leakage |\n| 0.05 -- 0.10 | Good | Typical for strong factors |\n| 0.02 -- 0.05 | Moderate | Profitable with enough breadth |\n| < 0.02 | Weak | Needs very high capacity to matter |\n\n## The Fundamental Law of Active Management\n\n$$IR \\approx IC \\times \\sqrt{BR}$$\n\nWhere IR is the information ratio and BR is breadth (independent bets per year). IC = 0.03 across 500 stocks rebalanced monthly: IR \u2248 0.03 \u00d7 \u221a6000 \u2248 2.3. Low IC is profitable with enough breadth.\n\n## Overlap Inflation\n\nOverlapping return labels (e.g., 21-day forward returns sampled daily) reduce effective sample size to ~N/H and inflate IC. A \"significant\" IC on 1,000 daily observations with H=21 overlap has only ~48 independent points. Always report IC_IR from non-overlapping periods or adjust standard errors for overlap.\n\n## Guardrails\n\n- Always use Rank IC (Spearman) for cross-sectional signals -- Pearson is sensitive to outliers.\n- Report IC_IR (mean/std), not just mean IC -- consistency matters more than magnitude.\n- Use HAC (Newey-West) standard errors, not naive t-tests -- IC series are autocorrelated.\n- Plot IC over time: a decaying IC trend means the signal is crowding or the regime has changed.\n\n## Production Implementation\n\n`ml4t-diagnostic` provides IC computation with proper statistical testing:\n\n```python\nfrom ml4t.diagnostic.api import compute_ic_hac_stats, compute_ic_series\n\nic = compute_ic_series(\n predictions,\n returns,\n pred_col=\"prediction\",\n ret_col=\"forward_return\",\n date_col=\"date\",\n entity_col=\"symbol\",\n method=\"spearman\",\n)\nstats = compute_ic_hac_stats(ic, maxlags=5)\n# stats[\"mean_ic\"], stats[\"t_stat\"], stats[\"p_value\"]\n```\n\n## Checklist\n\n- [ ] IC computed per cross-section (not pooled across all dates)\n- [ ] Rank IC (Spearman) used for cross-sectional signals\n- [ ] IC_IR reported alongside mean IC\n- [ ] Statistical significance tested with HAC standard errors (not naive t-test)\n- [ ] IC decay by horizon plotted to confirm signal persistence\n"
}
},
{
"turn": 4,
"type": "tool_call",
"name": "read_skill",
"args": {
"name_or_path": "~/ml4t/skills/infrastructure/registry-system/SKILL.md"
},
"result": {
"path": "~/ml4t/skills/infrastructure/registry-system/SKILL.md",
"size_bytes": 4880,
"content": "---\nname: ml4t-registry-system\ndescription: \"Content-addressed experiment tracking for ML trading models. Use when versioning models, features, or experiment artifacts for reproducibility.\"\nwhen_to_use: \"Use when running model experiments and need reproducibility, comparison, and audit trail across training runs\"\ndependencies: []\nmetadata:\n book_chapters: \"11, 12\"\n library: \"\"\npaths: [\"**/*schema*.py\", \"**/*registry*.py\", \"**/*pipeline*.py\", \"**/*polars*.py\", \"**/*case_study*.py\"]\n---\n# Experiment Registry\n\nWithout a registry, you overwrite the best model every time you retrain. Content-addressed storage \u2014 where hash(config) determines the storage path \u2014 makes every experiment reproducible and comparable without manual bookkeeping.\n\n## The Problem\n\nA quant runs 50 model configurations. Results go into `model_v2_final_FINAL.pkl`. Next week, a new run overwrites it. The team cannot answer: which hyperparameters produced the best IC? Was that before or after the feature change? Did we already try alpha=0.01? Without structured tracking, experiments are lost, repeated, and unverifiable.\n\n## The Pattern\n\n### WRONG\n```python\nimport pickle\n\n# Overwrite on every run \u2014 no history, no comparison, no provenance\nmodel.fit(X_train, y_train)\nwith open(\"best_model.pkl\", \"wb\") as f:\n pickle.dump(model, f)\n\n# Three weeks later: \"Which config was this? What data did it use?\"\n```\n\n### CORRECT\n```python\nimport hashlib\nimport json\nimport sqlite3\nfrom datetime import datetime\nfrom pathlib import Path\n\ndef config_hash(config: dict) -> str:\n \"\"\"Deterministic hash of experiment config.\"\"\"\n blob = json.dumps(config, sort_keys=True).encode()\n return hashlib.sha256(blob).hexdigest()[:12]\n\ndef register_run(db_path: str, config: dict, metrics: dict, predictions_path: str):\n \"\"\"Register a training run with full provenance.\"\"\"\n run_hash = config_hash(config)\n conn = sqlite3.connect(db_path)\n conn.execute(\"\"\"\n CREATE TABLE IF NOT EXISTS training_runs (\n run_hash TEXT PRIMARY KEY,\n config JSON NOT NULL,\n metrics JSON NOT NULL,\n predictions_path TEXT,\n created_at TEXT NOT NULL\n )\n \"\"\")\n conn.execute(\n \"INSERT OR REPLACE INTO training_runs VALUES (?, ?, ?, ?, ?)\",\n (run_hash, json.dumps(config), json.dumps(metrics),\n predictions_path, datetime.now().isoformat()),\n )\n conn.commit()\n return run_hash\n\n# Usage: every config gets a unique, reproducible slot\nconfig = {\"model\": \"ridge\", \"alpha\": 1.0, \"features\": \"momentum_v2\"}\nrun_hash = register_run(\"registry.db\", config, {\"ic\": 0.04}, f\"runs/{config_hash(config)}/predictions.parquet\")\n# Re-running same config overwrites same slot \u2014 idempotent\n```\n\n## Registry Schema\n\nThree linked tables capture the full experiment lifecycle:\n\n```\ntraining_runs prediction_sets backtest_runs\n+------------+ +----------------+ +--------------+\n| run_hash |<------>| pred_hash |<----->| bt_hash |\n| config | 1:N | run_hash (FK) | 1:N | pred_hash(FK)|\n| metrics | | fold | | config |\n| created_at | | path | | metrics |\n+------------+ +----------------+ +--------------+\n```\n\n- **training_runs**: one row per unique model config (hash of hyperparams)\n- **prediction_sets**: one row per fold or time split within a training run\n- **backtest_runs**: one row per strategy config applied to a prediction set\n\n## Content-Addressed Storage\n\n```\nrun_log/\n registry.db # SQLite: all metadata\n models/{config_hash}/ # hash(config) -> directory\n config.json\n metrics.json\n predictions.parquet\n```\n\nThe hash is the directory name. Same config always maps to the same directory. No manual naming, no collisions, no \"v2_final\" suffixes. Query the registry with standard SQL against `registry.db`.\n\n## Guardrails\n\n- Hash must be deterministic: `json.dumps(config, sort_keys=True)` \u2014 without `sort_keys`, same config produces different hashes\n- Register per-config as they complete, not in bulk after all finish \u2014 a crash at config 49 of 50 loses everything otherwise\n- Never store model weights in the SQLite database \u2014 store paths to artifacts on disk\n- Config must capture everything needed to reproduce: model type, hyperparameters, feature version, data version, random seed\n- Old runs are never deleted \u2014 mark as superseded, keep for audit trail\n\n## Checklist\n\n- [ ] Every experiment has a deterministic config hash\n- [ ] Registry stores config, metrics, and artifact paths (not weights in DB)\n- [ ] Runs registered incrementally (per-config, not bulk)\n- [ ] Same config re-run maps to same hash (idempotent)\n- [ ] Top-N query by any metric works against the registry\n- [ ] Full provenance: model type, hyperparams, feature version, data version, seed\n"
}
},
{
"turn": 4,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT * FROM training_runs LIMIT 1"
},
"result": {
"columns": [
"training_hash",
"family",
"label",
"config_name",
"spec_json",
"created_at",
"git_commit",
"entry_point",
"started_at",
"elapsed_s"
],
"rows": [
{
"training_hash": "1957228a8dcb",
"family": "linear",
"label": "fwd_ret_21d",
"config_name": "ridge_a1000000.0",
"spec_json": "{\"config_name\":\"ridge_a1000000.0\",\"family\":\"linear\",\"feature_sets\":[\"financial\",\"model_based\"],\"label\":\"fwd_ret_21d\",\"library\":\"sklearn\",\"n_folds\":8,\"params\":{\"alpha\":1000000.0},\"seed\":42}",
"created_at": "2026-04-10T11:16:46.370935+00:00",
"git_commit": "8c6eda7",
"entry_point": "06_linear",
"started_at": "2026-04-10T11:16:31.883058+00:00",
"elapsed_s": 4.310523570049554
}
],
"n_rows_returned": 1,
"truncated_at_200": false
}
},
{
"turn": 5,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT t.family, t.config_name, t.label, p.prediction_hash, p.split, p.training_hash\nFROM prediction_sets p\nJOIN training_runs t ON p.training_hash = t.training_hash\nWHERE t.label = 'fwd_ret_21d'\nORDER BY t.family, t.config_name, p.split"
},
"result": {
"columns": [
"family",
"config_name",
"label",
"prediction_hash",
"split",
"training_hash"
],
"rows": [
{
"family": "deep_learning",
"config_name": "lstm_h64",
"label": "fwd_ret_21d",
"prediction_hash": "3bc3fbdcc43d",
"split": "holdout",
"training_hash": "0488120b490e"
},
{
"family": "deep_learning",
"config_name": "lstm_h64",
"label": "fwd_ret_21d",
"prediction_hash": "57b4d1ffd748",
"split": "validation",
"training_hash": "0488120b490e"
},
{
"family": "deep_learning",
"config_name": "lstm_h64",
"label": "fwd_ret_21d",
"prediction_hash": "affdaa459d15",
"split": "validation",
"training_hash": "0488120b490e"
},
{
"family": "deep_learning",
"config_name": "nlinear",
"label": "fwd_ret_21d",
"prediction_hash": "820be2507261",
"split": "validation",
"training_hash": "a0d38ddfa83a"
},
{
"family": "deep_learning",
"config_name": "tsmixer",
"label": "fwd_ret_21d",
"prediction_hash": "dd35f0070431",
"split": "validation",
"training_hash": "c8af24d6fe4e"
},
{
"family": "gbm",
"config_name": "default_huber",
"label": "fwd_ret_21d",
"prediction_hash": "0c0035b1c7c9",
"split": "validation",
"training_hash": "a5d3b55cc482"
},
{
"family": "gbm",
"config_name": "default_mae",
"label": "fwd_ret_21d",
"prediction_hash": "dd021a11fd50",
"split": "validation",
"training_hash": "9b8d0d62422c"
},
{
"family": "gbm",
"config_name": "default_mse",
"label": "fwd_ret_21d",
"prediction_hash": "e2cb846c3d12",
"split": "validation",
"training_hash": "ab10713af672"
},
{
"family": "gbm",
"config_name": "leaves_15_huber",
"label": "fwd_ret_21d",
"prediction_hash": "1ea4b9261805",
"split": "validation",
"training_hash": "3040ebdc3ea4"
},
{
"family": "gbm",
"config_name": "leaves_15_mae",
"label": "fwd_ret_21d",
"prediction_hash": "9d5116906107",
"split": "validation",
"training_hash": "e595f259f0d0"
},
{
"family": "gbm",
"config_name": "leaves_15_mse",
"label": "fwd_ret_21d",
"prediction_hash": "3ab37612c798",
"split": "validation",
"training_hash": "d6addf7a3354"
},
{
"family": "gbm",
"config_name": "leaves_31_huber",
"label": "fwd_ret_21d",
"prediction_hash": "42c693540443",
"split": "validation",
"training_hash": "45aa68e23da9"
},
{
"family": "gbm",
"config_name": "leaves_31_mae",
"label": "fwd_ret_21d",
"prediction_hash": "128e4d7cbdaf",
"split": "validation",
"training_hash": "6855c1f66909"
},
{
"family": "gbm",
"config_name": "leaves_31_mse",
"label": "fwd_ret_21d",
"prediction_hash": "8fa4854ae3f6",
"split": "validation",
"training_hash": "7c9cd3775519"
},
{
"family": "gbm",
"config_name": "leaves_63_huber",
"label": "fwd_ret_21d",
"prediction_hash": "61b85b361b7d",
"split": "validation",
"training_hash": "fcba796cd9a0"
},
{
"family": "gbm",
"config_name": "leaves_63_mae",
"label": "fwd_ret_21d",
"prediction_hash": "507e5745de47",
"split": "validation",
"training_hash": "3bf78b212b1f"
},
{
"family": "gbm",
"config_name": "leaves_63_mse",
"label": "fwd_ret_21d",
"prediction_hash": "bf6c049c4751",
"split": "validation",
"training_hash": "71ca6dae76cd"
},
{
"family": "gbm",
"config_name": "leaves_7_huber",
"label": "fwd_ret_21d",
"prediction_hash": "8748613e0c9a",
"split": "validation",
"training_hash": "cd8f0c4c2235"
},
{
"family": "gbm",
"config_name": "leaves_7_mae",
"label": "fwd_ret_21d",
"prediction_hash": "226acb97bed9",
"split": "validation",
"training_hash": "d259643d2a2d"
},
{
"family": "gbm",
"config_name": "leaves_7_mse",
"label": "fwd_ret_21d",
"prediction_hash": "e716461a3d25",
"split": "validation",
"training_hash": "b4dbc2e98a77"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "b7efee934c14",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "57d6fc10d765",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "b39853c9a87e",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "a781d0703176",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "b5e8074e885f",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "f54c1b723181",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "5650980967e0",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "3b6353663abe",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "0030adc1655e",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "cae",
"label": "fwd_ret_21d",
"prediction_hash": "be628911cbd7",
"split": "validation",
"training_hash": "3e770b25226e"
},
{
"family": "latent_factors",
"config_name": "ipca",
"label": "fwd_ret_21d",
"prediction_hash": "354ea2bf6afc",
"split": "validation",
"training_hash": "79f4fed60254"
},
{
"family": "latent_factors",
"config_name": "pca",
"label": "fwd_ret_21d",
"prediction_hash": "05fb5bcf7ca1",
"split": "validation",
"training_hash": "2e13f7232f9f"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "d3248c933af2",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "0ab1b2818259",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "9109d08cff6a",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "01f1f2587418",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "bed43a0a327d",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "77ce5dfe2575",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "b690d1bba817",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "e5d5883fab3a",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "48fa2a1e92f1",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sae",
"label": "fwd_ret_21d",
"prediction_hash": "9b0a631011a3",
"split": "validation",
"training_hash": "e0c6034e33d0"
},
{
"family": "latent_factors",
"config_name": "sdf",
"label": "fwd_ret_21d",
"prediction_hash": "98ecb12aa7b9",
"split": "validation",
"training_hash": "9c9e121dea66"
},
{
"family": "latent_factors",
"config_name": "sdf",
"label": "fwd_ret_21d",
"prediction_hash": "b7da3311263f",
"split": "validation",
"training_hash": "9c9e121dea66"
},
{
"family": "latent_factors",
"config_name": "sdf",
"label": "fwd_ret_21d",
"prediction_hash": "2a0b20019c21",
"split": "validation",
"training_hash": "9c9e121dea66"
},
{
"family": "latent_factors",
"config_name": "sdf",
"label": "fwd_ret_21d",
"prediction_hash": "feaa7395fe89",
"split": "validation",
"training_hash": "9c9e121dea66"
},
{
"family": "latent_factors",
"config_name": "sdf",
"label": "fwd_ret_21d",
"prediction_hash": "7801dd3ed352",
"split": "validation",
"training_hash": "9c9e121dea66"
},
{
"family": "linear",
"config_name": "enet_a0.01",
"label": "fwd_ret_21d",
"prediction_hash": "962f7289cb3f",
"split": "validation",
"training_hash": "9b433bc614a4"
},
{
"family": "linear",
"config_name": "ols",
"label": "fwd_ret_21d",
"prediction_hash": "b7cae60e1a41",
"split": "validation",
"training_hash": "31ac8f2c4096"
},
{
"family": "linear",
"config_name": "ridge_a0.001",
"label": "fwd_ret_21d",
"prediction_hash": "cd23437a2f14",
"split": "validation",
"training_hash": "fc6a2c5dda8d"
},
{
"family": "linear",
"config_name": "ridge_a0.01",
"label": "fwd_ret_21d",
"prediction_hash": "1d35ff9b5a05",
"split": "validation",
"training_hash": "2617a804b1f4"
},
{
"family": "linear",
"config_name": "ridge_a0.1",
"label": "fwd_ret_21d",
"prediction_hash": "5ba93088d3d1",
"split": "validation",
"training_hash": "6f3de862b64e"
},
{
"family": "linear",
"config_name": "ridge_a1.0",
"label": "fwd_ret_21d",
"prediction_hash": "521b0b81b305",
"split": "validation",
"training_hash": "31529f2d0f78"
},
{
"family": "linear",
"config_name": "ridge_a10.0",
"label": "fwd_ret_21d",
"prediction_hash": "6deb5ff14199",
"split": "validation",
"training_hash": "6781dd2af9b8"
},
{
"family": "linear",
"config_name": "ridge_a100.0",
"label": "fwd_ret_21d",
"prediction_hash": "2855922eab98",
"split": "validation",
"training_hash": "697afbbdf17f"
},
{
"family": "linear",
"config_name": "ridge_a1000.0",
"label": "fwd_ret_21d",
"prediction_hash": "086fa8ab2dc0",
"split": "validation",
"training_hash": "7226f3f2d844"
},
{
"family": "linear",
"config_name": "ridge_a10000.0",
"label": "fwd_ret_21d",
"prediction_hash": "b91b8d68552c",
"split": "validation",
"training_hash": "ecbb7e087782"
},
{
"family": "linear",
"config_name": "ridge_a100000.0",
"label": "fwd_ret_21d",
"prediction_hash": "1f0d72dc1880",
"split": "validation",
"training_hash": "42e86ed5e992"
},
{
"family": "linear",
"config_name": "ridge_a1000000.0",
"label": "fwd_ret_21d",
"prediction_hash": "d533746293bd",
"split": "validation",
"training_hash": "1957228a8dcb"
},
{
"family": "linear",
"config_name": "ridge_a10000000.0",
"label": "fwd_ret_21d",
"prediction_hash": "8f9f5f01bf9e",
"split": "validation",
"training_hash": "319a2f444296"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "99359adfad70",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "e039d5622929",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "eb8638783b2a",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "51aaab130816",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "5d2793caff56",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "97770b304e55",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "445e393470c3",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_l",
"label": "fwd_ret_21d",
"prediction_hash": "b5ce78f61750",
"split": "validation",
"training_hash": "d5baaadd0c54"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "b172aa6b9db2",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "f67790943264",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "c4c1ab59fc7f",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "cdeaeca48ea7",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "1a540e1c12ac",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "3e50c0f98555",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "fa3a671a048d",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_m",
"label": "fwd_ret_21d",
"prediction_hash": "25461e7ee8e5",
"split": "validation",
"training_hash": "e0d0889cbe48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "10a2d7337609",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "1b5937d62cb1",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "76325d7460b7",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "53cfeb4afc25",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "e90c50265bb5",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "d0cb802fd9d6",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "e1db89148d91",
"split": "validation",
"training_hash": "15a6ab4f5e48"
},
{
"family": "tabular_dl",
"config_name": "tabm_s",
"label": "fwd_ret_21d",
"prediction_hash": "b8f6a9781c1b",
"split": "validation",
"training_hash": "15a6ab4f5e48"
}
],
"n_rows_returned": 84,
"truncated_at_200": false
}
},
{
"turn": 6,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT * FROM prediction_metrics LIMIT 1"
},
"result": {
"columns": [
"prediction_hash",
"computed_at",
"ic_mean",
"ic_std",
"ic_t",
"n_folds",
"pct_positive",
"task_type",
"accuracy",
"balanced_accuracy",
"auc_roc",
"auc_pr",
"log_loss",
"brier_score",
"ic_mean_daily",
"ic_std_daily",
"ic_n_days",
"ic_pct_positive",
"ic_se_naive",
"ic_naive_lo",
"ic_naive_hi",
"ic_se_hac",
"ic_ci_lo",
"ic_ci_hi",
"ic_t_hac",
"ic_p_hac",
"ic_hac_lag",
"ic_boot_lo",
"ic_boot_hi",
"ic_boot_block",
"n_obs"
],
"rows": [
{
"prediction_hash": "d533746293bd",
"computed_at": "2026-05-02T18:32:13.905593+00:00",
"ic_mean": 0.05354295687257081,
"ic_std": 0.0402747869735435,
"ic_t": 3.760227252275156,
"n_folds": 8.0,
"pct_positive": 1.0,
"task_type": "regression",
"accuracy": null,
"balanced_accuracy": null,
"auc_roc": null,
"auc_pr": null,
"log_loss": null,
"brier_score": null,
"ic_mean_daily": 0.05354295687257081,
"ic_std_daily": 0.31765471385635646,
"ic_n_days": 2016.0,
"ic_pct_positive": 0.5654761904761905,
"ic_se_naive": 0.007074732777048802,
"ic_naive_lo": 0.03966840138495705,
"ic_naive_hi": 0.06741751236018458,
"ic_se_hac": 0.02247369860709252,
"ic_ci_lo": 0.00946884295726267,
"ic_ci_hi": 0.09761707078787896,
"ic_t_hac": 2.382471964613474,
"ic_p_hac": 0.01728909572974624,
"ic_hac_lag": 20.0,
"ic_boot_lo": 0.01428242197760413,
"ic_boot_hi": 0.09371599638639924,
"ic_boot_block": 21.0,
"n_obs": null
}
],
"n_rows_returned": 1,
"truncated_at_200": false
}
},
{
"turn": 6,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT * FROM fold_metrics LIMIT 1"
},
"result": {
"columns": [
"prediction_hash",
"fold_id",
"computed_at",
"ic",
"ic_std",
"n_entities",
"rmse",
"mae",
"accuracy",
"balanced_accuracy",
"auc_roc",
"auc_pr",
"log_loss",
"brier_score",
"auc_class_-1",
"auc_class_0",
"auc_class_1"
],
"rows": [
{
"prediction_hash": "d533746293bd",
"fold_id": 0,
"computed_at": "2026-04-10T11:16:46.585860+00:00",
"ic": 0.01890712262670275,
"ic_std": 0.3162404855366452,
"n_entities": 96.0,
"rmse": 0.05144636751619291,
"mae": 0.03768936624915362,
"accuracy": null,
"balanced_accuracy": null,
"auc_roc": null,
"auc_pr": null,
"log_loss": null,
"brier_score": null,
"auc_class_-1": null,
"auc_class_0": null,
"auc_class_1": null
}
],
"n_rows_returned": 1,
"truncated_at_200": false
}
},
{
"turn": 7,
"type": "tool_call",
"name": "query_registry",
"args": {
"sql": "SELECT t.family, t.config_name, p.prediction_hash, p.split, m.ic_mean, m.ic_t_hac, m.ic_p_hac, m.ic_ci_lo, m.ic_ci_hi\nFROM prediction_metrics m\nJOIN prediction_sets p ON m.prediction_hash = p.prediction_hash\nJOIN training_runs t ON p.training_hash = t.training_hash\nWHERE t.label = 'fwd_ret_21d' AND p.split = 'validation'\n AND t.family IN ('gbm', 'tabular_dl', 'latent_factors')\nORDER BY t.family, m.ic_mean DESC"
},
"result": {
"columns": [
"family",
"config_name",
"prediction_hash",
"split",
"ic_mean",
"ic_t_hac",
"ic_p_hac",
"ic_ci_lo",
"ic_ci_hi"
],
"rows": [
{
"family": "gbm",
"config_name": "leaves_7_mse",
"prediction_hash": "e716461a3d25",
"split": "validation",
"ic_mean": 0.03712574068456292,
"ic_t_hac": 1.8816464785608475,
"ic_p_hac": 0.06002806819108786,
"ic_ci_lo": -0.0015684825083625287,
"ic_ci_hi": 0.07581996387748838
},
{
"family": "gbm",
"config_name": "lea出典を明記したうえで、ライセンスに従って全文を掲載しています。 ライセンス: MIT
この要約は原文をもとにStratmillのリサーチエージェントが作成したもので、出典の複製ではありません。