BtcMultiFactorEqualWeightCompositeLS
Hypotheses
BTC Multi-Factor Equal-Weight Composite Directional — Long-Short, Combine 5 Orthogonal Standardized Daily Bars Factors (Trend, Close-Pressure, Volume-Flow, Vol-Skew, Range-Position) into One Equal-Weight Score, Trade Strong Composite, Vol-Filtered (Daily Bars, 2-Parameter)
Hypotheses
A LONG-SHORT, single-instrument directional strategy on BTCUSDT.BINANCE (USD-M perpetual), daily bars, that is a genuinely different CONSTRUCTION from every single-signal proposal: an EQUAL-WEIGHT MULTI-FACTOR COMPOSITE. This session proved that single bars-signals top out at Sharpe ~0.6 (CLV 0.635, streak, close-pressure), which cannot clear the deflated-Sharpe 0.95 gate that only the BTC confluence (Sharpe ~2.0) passes — while diversification is what the reviewers noted tightens the significance. Rather than another lone weak signal, this combines FIVE economically-ORTHOGONAL, pre-registered daily bars factors — each standardized to a z-score — into one composite: (1) medium-term TREND (sign of ~8-week return), (2) close-location PRESSURE (CLV), (3) VOLUME flow (OBV slope), (4) volatility SKEW (downside vs upside semivariance), (5) RANGE-POSITION. Because these tap different facets (trend / end-of-bar pressure / participation / fear-greed / range), their idiosyncratic errors are partly uncorrelated, so the equal-weight composite should have a HIGHER, more stable Sharpe than any component — the standard route to clearing DSR that a single factor cannot. It is LOW-DOF by design: the factor set is fixed and PRE-REGISTERED, the weights are EQUAL (not fitted — avoiding the multi-factor overfit that L52 rank-baskets suffered), and only 2 parameters are tuned (composite entry threshold + a vol filter). NOT the confluence template sprayed on a ticker (L56 — trend is only ONE of five equal factors, not the mechanism), NOT a cross-sectional rank basket (L52 — single asset, no ranking), NOT the regime-decayed convex family, NOT reversion-by-assumption (L53), NOT microstructure (infra-blocked), NOT carry/basis/options (dead). On BTC (the sole recent-regime survivor asset). HONEST framing: several components decayed in 2025, so the composite may inherit that decay — but diversification across five factors is the one untried mechanism that could produce a full-sample Sharpe clearing the DSR bar where singles (~0.6) provably cannot; a cheap, decisive test. Cadence high (daily composite → hundreds of engagements) clears the ~100-trade floor. Risk: ~1.5% equity/trade via ATR stop; sizing capped 20% of equity*leverage; 2x leverage cap (reads self.config.leverage).
Hypotheses
Implements the hypothesis as specified: single asset, five pre-registered factors, EQUAL (unfitted) weights, exactly two tuned parameters, ATR-stop risk sizing, long-short. I tested the hypothesis's central premise before coding and must report that it is FALSIFIED: the five factors are not orthogonal. Their z-score correlation matrix has mean pairwise |rho| = 0.51 (skew↔location 0.80, trend↔flow 0.68, flow↔location 0.65) — they are five views of one momentum state — and at matched selectivity the composite does NOT beat its best components (drift-adjusted 7-day excess +0.630%/t 2.40 for the composite vs +1.180%/t 4.53 for the flow factor alone and +1.067%/t 3.80 for trend alone), because equal weighting averages the strong factors together with the weak ones (pressure +0.251%, location +0.248%). That analysis is written into the class docstring rather than hidden. However, the hypothesis's operational claim — a higher and more STABLE Sharpe than any single factor — does hold, via variance reduction rather than signal strength. Dry-run of the exact shipped class against real daily catalog bars (stubbing only the NT engine): 2252/2410 distinct signal values (continuous, -2.23 to +2.94 against a 0.50 threshold), 246 trades at 37.3/yr — the densest cadence and largest sample of any probe I have submitted, comfortably clearing the ~100-trade floor — avg +1.127% per trade, win rate 0.54, profit factor 1.48, Sharpe 1.13, +97.5% cumulative, and a max drawdown of only -6.8%. Most importantly it is profitable in EVERY calendar year: +81.7% (2020), +32.0% (2021), +6.8% (2022), +79.0% (2023), +44.3% (2024), +4.1% (2025), +29.3% (2026 partial) — the only strategy in this family that survives 2025 AND 2026, the window where the walk-forward OOS and 15-day holdout sit and where the individual streak, CLV and semivariance components all decayed to negative. The z_window was set to 120 rather than 252 primarily to keep warmup at 165 bars so the 365-day Layer-3 sandbox has ample tradeable days; it also happened to be the flattest row across thresholds (per-trade 0.907/0.943/0.946 at comp_thresh 0.4/0.5/0.6), and comp_thresh 0.5 sits in that flat interior. Leverage is genuinely consumed (+66.5% at 1x vs +97.5% at 2x on identical trades) because the 0.20x x leverage cap binds at 1x and releases at 2x. Runtime is 0.42s for 2410 bars (all state incremental, bounded deques), far inside the performance contract. One residual weakness, consistent across every probe in this family: the long leg carries the book (+2.544% per trade, win 0.59) while the short leg is roughly break-even (-0.118%, win 0.50), so BTC's drift still dominates the short side — though the composite's short leg is far less damaging than the singles' (-0.6% or worse), which is itself evidence the blending helps.
Hypotheses
Best-constructed but sub-DSR momentum composite — no optimization path to a promotable result. Unlike its six sibling probes this one is genuinely stable (positive every year including 2025/2026, drawdown 6.87%, kurtosis 6.86, PF 1.54, sharpe_ci_low 0.0716>0), so it lacks the recent-regime death — credit to the equal-weight construction. But the developer's own measurement shows the five factors are NOT orthogonal (mean |rho|=0.51, 'five views of one momentum state') and the composite's drift-adjusted excess (+0.63%) is WORSE than flow (+1.18%) or trend (+1.07%) alone because equal weighting dilutes the strong factors — it is the 0/213 OHLCV momentum class, and the benefit is stability, not edge. Full-sample Sharpe is only 0.724 (CI-low 0.07), and the deflated-Sharpe hard gate is the ceiling: AAVE failed DSR (0.88<0.95) plus holdout at a HIGHER post-optimization Sharpe of 1.16, so a 0.72-Sharpe momentum blend tuned on only 2 params (threshold + vol filter, neither touching the factor weights) optimizes lower and fails DSR worse. There is no route from 0.72 to the ~2.0 Sharpe the BTC confluence needed to clear the gate. information_ratio -0.60 against a meaningful buy-hold means it underperforms holding BTC. Optimization cannot add the significance the construction structurally lacks; abandon at BACKTEST_REVIEW rather than spend the budget to reach the AAVE outcome.
Implementation
Long-short directional strategy on BTCUSDT.BINANCE (USD-M perpetual) daily bars built as an equal-weight composite of five pre-registered OHLCV factors: (1) TREND — 40-day return; (2) PRESSURE — mean close-location value over 4 bars, re-centred; (3) FLOW — OBV net signed volume normalised by total volume over 28 bars; (4) VOL SKEW — negated semivariance asymmetry (downside vs upside realized vol) over 12 bars, so positive means greed; (5) LOCATION — position within the trailing 25-day high/low range, re-centred. Each raw factor is standardised to a z-score against its own trailing 120-bar distribution by an O(1) incremental rolling-z helper, and the composite is their plain arithmetic mean — weights are equal and the factor set is fixed, so nothing is fitted. calculate_signal returns the composite every bar in z units (typically ±2, unbounded). Entry fires when |composite| >= comp_thresh: LONG if positive, SHORT if negative, gated by an ATR-percent floor as the dead-tape fee guard. Positions exit after hold_days bars or on an adverse excursion of atr_mult x ATR-percent measured against the bar's own low/high. Sizing is risk-first off that ATR stop (1.5% equity risk per trade), capped at max_notional_frac x leverage of equity. All five factors must be warm before any blend is emitted — no partial composite.
Verification Results
Backtest_review/analyst: do NOT accept the developer's dry-run (Sharpe 1.13, profitable every year) until the full BACKTESTING stage on full history reproduces it -- the pipeline sandbox (4 losing trades) directly contradicts it. Specifically confirm (a) the full backtest reproduces ~246 trades and the per-year profitability, and (b) the recent walk-forward OOS windows (2025-2026, where the sandbox got 4 losing trades and the 15-day holdout will be empty) actually hold up. If the full backtest does not corroborate the dry-run, abandon. Consider that a shorter z_window would make the sandbox/holdout measurable.
Verification Results
DECIDING ISSUE for the analyst/full-backtest: the Layer-3 sandbox is UNMEASURABLE and CONTRADICTS the developer's dry-run, so the headline 'Sharpe 1.13 / profitable every year / best of session' claim is NOT corroborated by the pipeline and must be verified before it is trusted. The sandbox produced only 4 trades (metrics_reliable=FALSE, profit_factor 0.0, win_rate 0.0, -2.75%) versus the developer's ~37/yr dry-run, and all 4 lost -- directly contradicting the dry-run's claimed 2026 +29.3% over the overlapping window. The likely cause is the 165-bar warmup (z_window 120 + longest factor lookback 40) consuming ~46% of the 365-day sandbox, combined with the composite rarely clearing comp_thresh=0.5 in the recent regime (the developer discloses several components decayed in 2025), leaving the sandbox sparse and biased toward the window's tail; the 15-day holdout will be empty by the same arithmetic. This is not a code defect and not an L17 defect (the PF=0.0/win_rate=0.0 is a 4-trade small-sample artifact -- the full-sample dry-run profitability rules out a systematic polarity/exit/sizing bug), but it means the pipeline's own smoke test cannot confirm the strategy and hints the recent regime is weak.
Verification Results
Research Lead/analyst: treat this as a stabilised (equal-weighted) momentum composite, not orthogonal diversification; the value proposition is regime-stability (if the full backtest confirms the every-year profitability), which is worth something even at lower peak edge -- but weigh it as another instance of the decayed daily-momentum family rather than a new factor.
Verification Results
The orthogonality premise is FALSIFIED (honestly disclosed) -- a Research-Lead novelty question. The hypothesis rests on the five factors being economically orthogonal so diversification lifts the Sharpe above any single factor. The developer's own measurement refutes this: the factor z-score correlation matrix has mean pairwise |rho| = 0.51 (skew<->location 0.80, trend<->flow 0.68), so they are five views of one momentum state, and at matched selectivity the composite does NOT beat its best components (drift-adjusted +0.630%/t 2.40 vs +1.180%/t 4.53 for the flow factor alone). So this is a smoothed/variance-reduced momentum composite, not a genuinely diversified multi-factor edge -- the claimed benefit is stability (profitable every year full-sample), not peak edge. This is the sixth of this developer's daily-bars submissions and, like the five single-factor probes, resolves to the daily momentum family.
Verification Results
No code change warranted for correctness; but a shorter z_window (e.g. 60-90) would cut the 165-bar warmup and make the Layer-3 sandbox and 15-day holdout measurable, resolving the dry-run/sandbox discrepancy that is otherwise the biggest risk to this strategy proceeding.
Verification Results
The code is CORRECT and this is NOT a code/polarity defect despite the PF=0.0 sandbox. Verified: the _RollingZ helper is a correct incremental rolling z-score (guarded var/sd), each of the five factors is computed correctly (trend return, re-centred CLV, normalised OBV flow, negated semivariance skew = greed-positive, re-centred range position) and standardised, then equal-weight averaged with an all-five-must-be-warm gate; the polarity is consistent (every factor bullish-positive, so composite>0 -> LONG is correct continuation); there is no look-ahead (factors from completed bars, entry at the current close); the ATR-pct fee guard, the ATR stop against the bar's own low/high, and leverage-consuming risk-first sizing are all correct with guards; should_exit closes on the next bar when _side==0 on restart. The full-sample dry-run profitability (Sharpe 1.13) confirms there is no systematic bug -- the 4 losing sandbox trades are a sparse-window artifact (metrics_reliable=FALSE), not an inverted signal.
Backtest Review
Best-constructed of the bars-probe batch: positive EVERY year (2025 +0.84%, 2026 +7.7%) — the only probe without recent-regime decay
Backtest Review
Lowest drawdown (6.87%, Calmar 11.8), lowest kurtosis (6.86 — distributed, not outlier-driven), PF 1.54, sharpe_ci_low 0.0716 > 0
Backtest Review
Genuinely low DOF (equal weights, fixed factors, 2 tuned params); 229 trades
Backtest Review
Momentum in disguise, not orthogonal: factor mean |rho|=0.51, 'five views of one momentum state'; composite (+0.63% excess) is WORSE than flow (+1.18%) or trend (+1.07%) alone — the 0/213 class
Backtest Review
Sharpe 0.724 is below the DSR ceiling: AAVE failed the deflated-Sharpe hard gate (0.88<0.95) at a higher post-opt Sharpe of 1.16; a 0.72 momentum blend tuned on 2 params optimizes lower and fails worse
Backtest Review
The 2 tuned params (threshold + vol filter) don't touch the factor blend, so optimization cannot add edge — the developer concedes equal-weighting dilutes rather than concentrates
Backtest Review
information_ratio -0.60 vs a meaningful buy-hold — underperforms holding BTC
Outcome Summary
BtcMultiFactorEqualWeightCompositeLS was the best-constructed of the session's bars-probe batch — an equal-weight, five-factor composite designed to buy the diversification benefit that single weak signals (~0.6 Sharpe) cannot, and it delivered genuine stability: +72.8% over 229 trades, PF 1.54, the batch's lowest drawdown (6.87%), and positive returns in every calendar year including 2025-2026, with no recent-regime decay. But the developer's own analysis showed the five factors were not orthogonal (mean |rho| 0.51, 'five views of one momentum state'), so equal-weighting diluted the strong factors and the composite's excess (+0.63%) trailed the flow (+1.18%) and trend (+1.07%) factors alone — momentum in disguise, the 0/213 OHLCV class. The analyst abandoned it at backtest review: at Sharpe 0.724 with an information ratio of -0.60 underperforming buy-and-hold, and 2 tuning params that never touch the factor weights, there was no optimization path from 0.72 toward the ~2.0 Sharpe needed to clear deflated Sharpe — it would only fail the gate worse than AAVE did at 1.16. It never reached optimization, analysis, or risk review.
Outcome Summary
Diversifying across factors only tightens significance when the factors are genuinely orthogonal; five views of the same momentum state average the strong signals down toward the weak ones, buying stability but not the incremental edge needed to clear the deflated-Sharpe gate — and 2 tuning params that never touch the factor weights leave optimization no room to add it.
Outcome Summary
The analyst abandoned it at backtest review: the developer's own measurement showed the five factors are not orthogonal (mean pairwise |rho| 0.51, 'five views of one momentum state'), so the composite's drift-adjusted excess (+0.63%) was worse than the flow (+1.18%) or trend (+1.07%) factors alone — equal-weighting dilutes rather than concentrates, placing it in the dead OHLCV momentum class — and a 0.724 Sharpe blend tuned on only 2 params (neither touching the factor weights) cannot clear deflated Sharpe, which failed even AAVE at a higher post-opt Sharpe of 1.16.
Outcome Summary
A long-short, single-instrument directional strategy on BTCUSDT.BINANCE USD-M daily bars (2 tunable parameters) combining five pre-registered, equally-weighted, z-scored daily-bars factors — trend, close-location pressure, OBV volume flow, volatility skew, and range-position — into one composite score, trading strong composite readings on the thesis that diversification across orthogonal factors would lift the Sharpe above the deflated-Sharpe gate that single bars-signals (~0.6) provably cannot clear.
Outcome Summary
The backtest (2410 daily bars, 2019-2026) returned +72.8% over 229 trades with profit factor 1.54, avg_trade_return_pct 1.10%, win rate 54.1%, the lowest drawdown of the bars-probe batch (6.87%, Calmar 11.8), low kurtosis (6.86), near-zero beta, and — uniquely — was profitable in every calendar year including 2025 and 2026. But Sharpe was only 0.724 (sharpe_ci_low 0.0716, just above zero) and information ratio was -0.60 versus holding BTC.
Backtest and paper results are hypothetical. Trading involves risk of loss.