BtcConvictionGatedMultiFactorLS
Hypotheses
BTC Conviction-Gated Multi-Factor Directional — Long-Short, Engage ONLY When a Diversified 5-Factor Composite AND the Medium-Term Trend Agree (Flat on Disagreement), Volatility-Targeted Sizing, Trailing-Stop Winners (Daily Bars, 3-Parameter)
Hypotheses
A LONG-SHORT, single-instrument directional strategy on BTCUSDT.BINANCE (USD-M perpetual), daily bars, that assembles the THREE ingredients this session's evidence identifies as separating the sole survivor (BTC confluence, Sharpe ~2.0) from the ~35 failed single-signal probes (all Sharpe ~0.55–0.64, DSR-failing): (1) DIVERSIFICATION — an equal-weight composite of orthogonal bars-factors (each single factor is only ~0.6 Sharpe, but combining uncorrelated ones raises the composite); (2) SELECTIVITY / flat-on-disagreement — the survivor trades ONLY when a fast and slow signal AGREE and stands flat otherwise, which is what produced its high, stable Sharpe; here the strategy engages ONLY when the composite AND the independent medium-term trend agree in direction; (3) VOLATILITY-TARGETING + trailing-stop winners — the survivor lets winners run on a trailing stop and its edge is risk-disciplined; here size is vol-targeted (downsizing in high vol) and winners are held on a trailing stop. It is NOT the confluence template sprayed on a ticker (L56 — the confirming signal is a 5-factor composite, not a second EMA; trend is one gate), NOT a single-signal probe (all Sharpe ~0.6), NOT the plain queued composite (this ADDS the selectivity gate + vol-targeting, the survival ingredients the plain composite lacks), NOT the regime-decayed convex family, NOT cross-sectional (L52 — single asset), NOT reversion (L53), NOT microstructure/carry/basis/options (dead/infra). LOW-DOF: the 5 factors are pre-registered and equal-weighted (not fitted), and only 3 parameters are tuned (composite threshold, trend lookback, trailing-ATR multiple). On BTC (the only recent-regime survivor asset). HONEST framing: the component factors decayed in 2025, so the composite may too; but the combination of diversification + selectivity + vol-targeting is the untried path most likely to lift a diversified bars edge over the DSR 0.95 bar that singles (~0.6) provably cannot clear — a decisive test of whether the survivor's edge is asset-specific or structure-specific. Risk: ~1.5% equity/trade anchored to the trailing stop; vol-targeted size capped 20% of equity*leverage; 2x leverage cap (reads self.config.leverage).
Hypotheses
Implements all three named ingredients on a single asset with exactly three tuned parameters (comp_thresh, trend_lookback, trail_atr_mult); the five factors, their lookbacks, the equal weights, z_window, init_atr_mult, max_hold_days and the vol-target settings are fixed by design. Dry-run of the exact shipped class against real daily catalog bars (stubbing only the NT engine): 2252/2410 distinct signal values (continuous, -2.23 to +2.94), 144 trades at 21.8/yr, avg +1.452% per trade (roughly 10x the 0.15% fee-viability floor), win rate 0.49, profit factor 1.48, Sharpe 0.90, +84.3% cumulative, max drawdown -10.7%, average hold 11.2 days, runtime 0.45s for 2410 bars with min_bars_required 165. Both legs are profitable (long +2.747% per trade over 70 trades, short +0.227% over 74) and the year-by-year record is positive in six of seven years including the recent regime that killed the single-factor probes: +75.4% (2020), +35.9% (2021), -1.5% (2022), +22.3% (2023), +30.3% (2024), +14.0% (2025), +32.6% (2026 partial). Leverage is genuinely consumed (+44.5% at 1x vs +84.3% at 2x) because sizing anchors to the 2-ATR initial stop rather than the wide trail, so the 0.20x x leverage cap binds at 1x and releases at 2x. I measured each ingredient in isolation before assembling, and must report that two of the three transfer and one does not. The trend-agreement gate helps (per-trade +0.943% to +1.093%, PF 1.37 to 1.45, drawdown -8.8% to -6.6%) and vol-targeting helps (Sharpe 1.02 to 1.09, cumulative +89.6% to +118.5%). The TRAILING STOP is the ingredient that does not transfer: the identical strategy with a plain fixed adverse-excursion stop and a 7-day hold scores Sharpe 1.04 at 1.7x the trade count and +106.3% cumulative, versus 0.85-0.90 and +68.6-84.3% here; pure unarmed trailing stops were worse still (Sharpe 0.45-0.61 across trail multiples 2.0-6.0, tested at three hold caps). I therefore implemented the best-performing form of the specified mechanism — a profit-armed trail that only ratchets after the trade has travelled the initial stop distance in its favour — and declared trail_atr_mult's bounds at [3.5, 5.5] where the trail stays wide enough not to cut winners on daily noise; performance is monotone improving in that multiple, so the optimizer will push toward the wide end. All 27 cells of the comp_thresh x trend_lookback x trail_atr_mult grid are profitable (Sharpe 0.41-1.08), and the shipped defaults sit in the interior rather than at the grid peak. If this iterates, replacing the trailing exit with the fixed stop plus 7-day hold is the single highest-value change and is worth roughly +0.15 Sharpe and a doubling of trade count.
Hypotheses
The hypothesis's own thesis — that selectivity + vol-targeting + trailing stops would lift the diversified composite over the DSR 0.95 bar — is FALSIFIED in this backtest: relative to the plain 5-factor composite it made things worse, Sharpe 0.724 -> 0.579 and sharpe_ci_low +0.072 -> -0.055 (now straddling zero), with trade count cut from 229 to 133. The developer further documents that the trailing-stop ingredient 'does NOT transfer to this signal' (a plain fixed stop scored Sharpe 1.04 vs 0.85), so a specified mechanism is conceded to degrade the result. The underlying signal is the same 0/213 OHLCV momentum composite (factors ~0.51 correlated), and a Sharpe of 0.579 with a CI on/below zero cannot clear the deflated-Sharpe hard gate after the optimizer's best-of-N — AAVE failed DSR (0.88<0.95) plus holdout at a much higher post-opt Sharpe of 1.16. information_ratio -0.615 against a meaningful buy-hold means it underperforms holding BTC. The 3 tuned params tune the degrading gate/trail machinery, not the factor edge, so optimization cannot recover what the construction gave away. This is the strictly-worse sibling of a composite already judged sub-DSR; abandon at BACKTEST_REVIEW rather than spend the optimization budget.
Implementation
Long-short directional strategy on BTCUSDT.BINANCE (USD-M perpetual) daily bars combining three mechanisms. DIVERSIFICATION: five pre-registered OHLCV factors — 40-day trend, 4-bar close-location pressure, 28-bar OBV net volume flow, 12-bar semivariance vol skew, and 25-day range position — each standardised to a z-score against its own trailing 120-bar distribution by an O(1) incremental rolling-z helper, then averaged with equal (unfitted) weights into a composite that calculate_signal returns every bar in z units. SELECTIVITY: the strategy engages only when the composite and an independent medium-term trend (sign of the trend_lookback-day return, a different lookback from the composite's internal 40-day trend factor) agree in direction, and stands flat on disagreement. VOL-TARGETING AND TRAILING WINNERS: position size scales by clamp(vol_target / realized_vol(20), 0.4, 1.6) so exposure shrinks when volatility runs hot; risk per trade is anchored to a fixed initial stop of init_atr_mult x ATR-percent, and once the trade's favourable excursion exceeds that same distance the stop converts to a trailing stop at trail_atr_mult x ATR-percent behind the running favourable extreme, so losers are cut at known risk while winners ratchet. A max_hold_days backstop closes anything still open. Sizing is risk-first off the initial stop, vol-scaled, then capped at max_notional_frac x leverage of equity.
Verification Results
Backtest_review/analyst: do NOT accept the dry-run (Sharpe 0.90, profitable every year) until the full BACKTESTING stage on full history reproduces ~144 trades AND the recent walk-forward OOS windows (2025-2026, where the sandbox got 2 losing trades and the holdout will be empty) hold up. The sandbox directly contradicts the claim. A shorter z_window would make the sandbox/holdout measurable.
Verification Results
DECIDING ISSUE: the Layer-3 sandbox is UNMEASURABLE and CONTRADICTS the dry-run -- even more so than the plain composite. The selectivity gate (composite AND trend must agree) plus the trailing stop cut cadence on top of the 165-bar warmup, so the sandbox produced only 2 trades (vs the plain composite's 4), both losing (-3.48%, metrics_reliable=FALSE, profit_factor 0.0, win_rate 0.0) -- directly contradicting the developer's dry-run claim of 2026 +32.6% and 'profitable 6/7 years including the recent regime'. This is NOT a code defect and NOT an L17 defect (the PF=0.0/win_rate=0.0 is a 2-trade small-sample artifact; the full-sample dry-run profitability, Sharpe 0.90, rules out a systematic polarity/exit/sizing bug), but the pipeline's own smoke test cannot confirm the strategy and the 15-day holdout will be empty by the same 165-bar-warmup arithmetic. The dry-run's positive recent regime is not corroborated by the pipeline and may be a warmup-selection artifact (profitable early-window trades eaten by warmup, losing later trades measured).
Verification Results
Research Lead/analyst: the highest-value next iteration (the developer's own measured recommendation) is to replace the trailing exit with the fixed stop + 7-day hold -- worth ~+0.15 Sharpe and a doubling of trade count, which would ALSO make the sandbox/holdout far more measurable (the low cadence here is partly the trailing stop's long holds). Treat the composite as a stabilised momentum sleeve, not orthogonal diversification.
Verification Results
Disclosed finding: the specified TRAILING STOP ingredient does not transfer, and the shipped construction is knowingly suboptimal -- a Research-Lead / next-iteration note, not a defect. The developer faithfully implemented the hypothesis's trailing-stop mechanism (correct behaviour: implement as specified, then report the measurement) and honestly reports it underperforms a plain fixed adverse-excursion stop + 7-day hold on every axis (Sharpe 1.04 vs 0.90, 2x the trade count, +106.3% vs +84.3% cumulative), with pure unarmed trailing stops worse still (Sharpe 0.45-0.61). So this is the best form of the SPECIFIED mechanism but not the best available exit. Separately, the composite's orthogonality premise is falsified (mean pairwise |rho| 0.51 -- it is a stabilised momentum composite, not diversification), as on the plain composite.
Verification Results
No code change warranted for correctness; but the fixed-stop swap the developer recommends both improves the edge and cures the sandbox/holdout sparsity, so it is the clear iteration if this proceeds.
Verification Results
The code is CORRECT, including the new armed-trailing-stop and vol-targeting -- this is NOT a code/polarity defect despite the PF=0.0 sandbox. Verified: the five-factor composite (rolling-z helper, equal-weight, bullish-positive polarity, composite>0 -> LONG) is the same verified structure as the plain composite; the selectivity gate (sign(composite)==sign(trend) else flat) is correct; the armed trailing stop is correct on both sides (peak tracks the favourable extreme, arms when the favourable excursion reaches init_dist, then trails at trail_atr_mult*ATR behind the peak, else holds the initial stop; risk-first sizing anchors to the INITIAL stop not the wide trail); vol-targeting scale = clamp(vol_target/realized_vol, 0.4, 1.6) is correct; there is no forward look-ahead (peak/stop use only the current bar's high/low, the standard daily trailing approximation); leverage is genuinely consumed. The full-sample dry-run profitability confirms no systematic bug -- the 2 losing sandbox trades are a sparse-window artifact (metrics_reliable=FALSE).
Backtest Review
PF 1.50, avg_trade_return_pct 1.43%, decorrelated (beta 0.006); positive most years incl. recent (2025 +1.2%, 2026 +16.4%)
Backtest Review
The added ingredients made it WORSE: vs the plain composite, Sharpe fell 0.724 -> 0.579 and sharpe_ci_low fell +0.072 -> -0.055 (now straddles zero); trade count cut 229 -> 133
Backtest Review
Developer concedes the trailing stop 'does NOT transfer' (plain fixed stop 1.04 vs 0.85 Sharpe) — a specified ingredient is documented as degrading
Backtest Review
Same 0/213 momentum composite base (factors ~0.51 correlated); Sharpe 0.579 cannot clear the DSR hard gate after best-of-N (AAVE failed DSR at a far higher 1.16)
Backtest Review
information_ratio -0.615 vs a meaningful buy-hold — underperforms holding BTC
Outcome Summary
BtcConvictionGatedMultiFactorLS was the capstone of the bars-probe session: it assembled the three ingredients that were thought to separate the sole BTC-confluence survivor (~2.0 Sharpe) from the ~35 failed single-signal probes — diversification, flat-on-disagreement selectivity, and vol-targeted sizing with trailing-stop winners — to test whether that edge was structure-specific. Instead it falsified its own thesis: relative to the plain 5-factor composite, Sharpe fell from 0.724 to 0.579, the confidence interval slid to straddle zero, and the developer's own ablation conceded the trailing stop degraded results (1.04 vs 0.85 Sharpe with a plain fixed stop). On the same ~0.51-correlated momentum composite base, a 0.579 Sharpe with an information ratio of -0.615 underperforming buy-and-hold cannot clear deflated Sharpe — a gate that failed even AAVE at 1.16 — and the 3 tuned params touched only the degrading machinery, not the factor edge. The analyst abandoned it at backtest review as the strictly-worse sibling of an already sub-DSR composite; it never reached optimization, analysis, or risk review.
Outcome Summary
The survivor's edge is asset/regime-specific, not a transferable recipe — grafting selectivity, vol-targeting, and trailing stops onto a diluted momentum composite subtracts rather than adds, and when a strategy's own ablation shows a specified ingredient (the trailing stop) degrades results and the whole is worse than its plainer sibling, there is no optimization path to recovery.
Outcome Summary
The analyst abandoned it at backtest review because the added ingredients made it strictly worse than the plain composite (Sharpe 0.724→0.579, sharpe_ci_low +0.072→-0.055, trades 229→133) — falsifying the hypothesis — and the developer's own measurement conceded the trailing stop 'does NOT transfer' (a plain fixed stop scored 1.04 vs 0.85 Sharpe); on the same ~0.51-correlated OHLCV momentum composite base, a 0.579 Sharpe with a CI on zero cannot clear deflated Sharpe (which failed even AAVE at 1.16), and the 3 params tune only the degrading gate/trail machinery, not the factor edge.
Outcome Summary
A long-short, single-instrument directional strategy on BTCUSDT.BINANCE USD-M daily bars (3 tunable parameters) that bolted the sole survivor's three ingredients — a diversified 5-factor equal-weight composite, a flat-on-disagreement selectivity gate (engage only when the composite and an independent medium-term trend agree), and volatility-targeted sizing with armed trailing-stop winners — onto the plain composite, testing whether the BTC confluence survivor's edge was structure-specific rather than asset-specific.
Outcome Summary
The backtest (2410 daily bars, 2019-2026) returned +58.6% over 133 trades with profit factor 1.50, avg_trade_return_pct 1.43%, near-zero beta, low drawdown (10.6%), and positive returns in most years including 2025 (+1.2%) and 2026 (+16.4%). But Sharpe fell to 0.579 with sharpe_ci_low -0.055 (CI now straddles zero) and information ratio was -0.615 versus holding BTC.
Backtest and paper results are hypothetical. Trading involves risk of loss.