Skip to content

View original

BinanceThreeMajorConfluenceTrendBasketLS

Hypotheses

Binance 3-Major Confluence Trend Basket, Long-Short (Proven 4H+1D Momentum-Confluence Signal Applied INDEPENDENTLY per Name to BTC/ETH/BNB USD-M Perps, Equal-Risk Legs, Net Exposure Floats, MULTI-YEAR History, 3 Shared Parameters)

Hypotheses

A LONG-SHORT, MULTI-INSTRUMENT, pure-OHLCV trend basket that runs this factory's ONLY optimization survivor — the 4H-primary + 1D-confirm momentum-confluence signal (BTC single-name version at Sharpe ~1.99, in paper) — INDEPENDENTLY on three decorrelated Binance USD-M majors (BTCUSDT, ETHUSDT, BNBUSDT) and holds the equal-risk portfolio average. The design is dictated by this session's failure analysis. The confluence signal itself works, but every attempt to prove NEW single-name variants keeps failing the multiple-testing gates (PBO > 0.5, DSR insignificant, holdout ratio < 0.70) because a single instrument's Sharpe is too noisy for a 225-trial optimizer to distinguish from best-of-N luck, and the Hyperliquid basket that would have diversified this died on HL's ~2.5-year history. This hypothesis fixes BOTH at once: (1) DIVERSIFICATION suppresses overfit — averaging three independent confluence legs cuts the idiosyncratic variance of the portfolio return, which directly lowers PBO and raises deflated Sharpe (a more stable objective is far harder to noise-select), the textbook antidote to the exact gate failures that killed the single names; (2) BINANCE MULTI-YEAR HISTORY (4+ years on all three names) delivers 300+ combined trades and long multi-regime walk-forward windows — clearing the ~100-trade measurability floor and giving the OOS enough data to hold up, the specific ingredient the HL basket and HL single-names lacked. It uses only 3 parameters SHARED identically across all three legs (no per-asset tuning), so parameter dimensionality is no higher than a single-name strategy while the edge is averaged over three assets. Asset choice maximizes decorrelation: BNB (exchange-token, burns, launchpad cycles) is the lowest-BTC-beta major, so its confluence timing differs from BTC/ETH, improving the aggregate Sharpe. It is materially different from every pipeline item: the single-name confluence ports (BTC/ETH/BNB) are correlated one-legged bets; the HyperliquidFiveMajorTrendFollowingBasket uses a dual-MA-state signal on HL's short/hostile history — THIS is the PROVEN confluence signal on Binance's multi-year history at the portfolio level. Fills long_short direction (13.6% vs target) and multi-instrument scope. Same venue, no cross-venue schema landmine. Pure OHLCV, no supplementary feed (avoids the short-history L/S-ratio/liquidations data_unavailable trap). I acknowledge it adds to the over-represented BINANCE venue; multi-year history — the proven survival ingredient — exists only there.

Hypotheses

I reused the validated confluence signal verbatim per leg and shared all three parameters across names, so parameter dimensionality matches a single-name strategy while the return stream averages three assets. The hypothesis's central claim is testable, so I tested it on real aligned data (14,164 bars, 2020-02-10 onward, 0.10% fees) rather than asserting it: single legs score Sharpe +0.31 (BTC), +0.52 (ETH), +0.52 (BNB), averaging +0.45, while the basket scores +0.55 on 2,359 combined trades. Crucially the leg-RETURN correlations — BTC-ETH 0.607, BTC-BNB 0.459, ETH-BNB 0.504 — sit well below the underlying asset correlations (0.84 and 0.72), because the confluence rule times each name independently, and the realized Sharpe gain matches the theoretical sqrt(3/(1+2*rho)) prediction at rho≈0.52 almost exactly. So the diversification mechanism is real, quantified, and behaves as the theory says it should, which is the property that should lower PBO. That measurement also caught a sizing trap I then fixed: at the single-name risk setting the three correlated legs stack to ~0.87x gross and their drawdowns compound, giving the basket a 42.4% drawdown against 19-25% for any single leg — near the abandon line and exactly the 'diversification made me feel safe, so I sized up' error. I cut per-leg risk from 1% to 0.45%, which leaves Sharpe unchanged (scaling is Sharpe-neutral) while bringing expected drawdown to roughly 25%, and I set the per-leg cap at 0.30 so it acts as a rare tail guard rather than an always-binding clamp that would render the risk sizing inert. On structure: all three legs share venue and timeframe, so the base class's alignment barrier defers the basket until every leg reports — which is what a portfolio decision needs — and it is safe because the three series share 14,164 timestamps with zero gaps; all legs including the primary are managed in one per-leg loop with the base hooks inert, so expect Layer 3's 'should_enter returned a side 0 times' diagnostic. Two caveats for the analyst. The trailing year is weak (+1.9%, Sharpe +0.24), consistent with all three single-name ports being flat to negative recently, so the +0.55 full-sample Sharpe is not evenly distributed. And +0.55 is modest in absolute terms: diversification here buys stability and a harder-to-noise-select objective, not magnitude — it makes a thin edge more robust, it does not make it large. One factual correction to the hypothesis's framing: BNB is described as the lowest-BTC-beta major, and by price correlation it is not (DOGE 0.551, XRP 0.620, SOL 0.651 all rank below BNB's 0.715), though BNB does happen to be the best diversifier of these three by leg correlation, so its inclusion is justified on the measurement that actually matters here even if the stated reason is wrong.

Hypotheses

No promotable edge, and the hypothesis's central diversification-as-robustness claim is empirically falsified by its own backtest. Over the largest possible sample (2148 trades, 6.6 years) the basket Sharpe is 0.319 — LOWER than the single-name ETH-Binance leg (0.43) already abandoned this session — because the three legs remained positively correlated (0.46-0.61) so drawdowns compounded (max_drawdown 27%, CI high 83.4%) instead of diversifying away. profit_factor 1.116 is < 1.2 (direct L9 abandon signal for OHLCV momentum, zero survivors), Sharpe CI [-0.29, 0.99] straddles zero, and information_ratio -0.539 shows negative risk-adjusted active value vs the benchmark. On top of the thin edge there is a fatal capacity blocker: impact_cost_pct 22.58% with capacity_usd only $1.96M — the edge is real only at toy scale, with market impact consuming ~a quarter of gross PnL ($41k on $100k). avg_trade_return_pct 0.459% clears the fee floor, so this is not a fee death, but the well-measured true edge (~0.32 Sharpe / 1.12 PF) is far below the 1.5 promotion floor and cannot be lifted by 3-parameter tuning without overfitting — any sub-window showing 1.5 would be a best-of-N selection that PBO/DSR reject, the exact death every sibling in this dual-TF-confluence family suffered. The mechanism is correctly coded and two-sided (no bug to iterate on). Abandon rather than spend 2 hours repeating the optimize→overfit-reject cycle on a diversification thesis the data already disproves.

Implementation

Long-short trend basket running the factory's validated dual-horizon momentum confluence signal independently on three Binance USD-M majors (BTCUSDT, ETHUSDT, BNBUSDT), 4H bars, with the equal-risk portfolio average held. Each leg goes long only when its fast ~2-day and slow ~7-day momentum are both positive, short when both are negative, and flat on disagreement; each exits on its own slow-leg flip or a 5x-ATR trailing stop. Net exposure floats freely from three legs long to three legs short. Each leg risks 0.45% of equity across its initial 1.5x-ATR stop with a 30% per-leg notional cap, giving roughly 0.53x gross when all three are aligned. Leverage 1.0. Exactly three tunable parameters shared identically across all legs (fast_lookback, slow_lookback, trail_atr_mult) — the same dimensionality as a single-name strategy — with ATR period, initial stop, per-leg risk, notional cap and per-symbol minimum notionals locked as constants.

Verification Results

Judge on cross-window walk-forward STABILITY and the deflated Sharpe / PBO (the metrics this design targets), not the headline total; confirm the edge is not solely the 2020-2021 window.

Verification Results

MODEST, REGIME-CONCENTRATED EDGE. The full-history basket Sharpe is +0.55 (+149%, 2,359 trades) — genuinely positive and measurable, but modest and not evenly distributed: the confluence legs all concentrate their edge in early years (2020-2021), and the trailing year is weak (developer: +1.9%, Sharpe +0.24; sandbox: -1.3%, Sharpe -0.04, avg_trade_return_pct +0.081% — positive but BELOW the 0.15% floor). Diversification buys robustness/stability, not magnitude.

Verification Results

Confirm at backtest_review that the full-history/optimized avg_trade_return_pct clears 0.15%; abandon only if the full sample stays below it.

Verification Results

RECENT PER-TRADE BELOW THE FEE FLOOR. Sandbox avg_trade_return_pct is +0.081%, under the 0.15% floor — but POSITIVE, and notably BETTER than the single-name trailing years (ETH -0.052%, BNB -0.281%), direct evidence the diversification smoothed the weak recent regime rather than a defect. The full-history per-leg edges (+0.39-0.43%) indicate the full-history basket per-trade should clear the floor; the sandbox reflects one weak year.

Verification Results

Accept BNB on the leg-correlation justification; the price-correlation premise is wrong but immaterial to the basket's diversification value. Note this adds to the over-represented Binance venue (acknowledged).

Verification Results

BNB INCLUSION — premise corrected by the developer. The hypothesis frames BNB as the 'lowest-BTC-beta major', which is factually wrong by price correlation (DOGE/XRP/SOL rank lower). BUT the developer correctly justifies BNB's inclusion on the metric that matters for a basket: its LEG-RETURN correlation to BTC (0.459) is the LOWEST of the three legs, so it is genuinely the best diversifier here. Unlike the BNB single-name I failed (lone 2021-artifact leg, decisively negative sandbox, no diversification value), BNB as 1/3 of a diversified, measurable basket with a flat sandbox is defensible.

Backtest Review

Huge decisive sample (2148 trades, 6.6 years), genuinely two-sided (1061 long / 1087 short); avg_trade_return_pct 0.459% clears the fee floor

Backtest Review

Correctly-implemented basket, low beta 0.025 + positive alpha 0.12; no code bug to iterate on

Backtest Review

Diversification thesis FALSIFIED: basket Sharpe 0.319 is LOWER than the single-name ETH leg (0.43) already abandoned — legs stayed correlated (0.46-0.61), drawdowns compounded (max_dd 27%, CI high 83.4%)

Backtest Review

profit_factor 1.116 < 1.2 — direct L9 abandon trigger for OHLCV momentum; Sharpe CI [-0.29, 0.99] straddles zero even over 2148 trades

Backtest Review

Capacity blocker: impact_cost_pct 22.58%, capacity_usd only $1.96M — the edge exists only at toy scale, market impact eats ~a quarter of gross PnL

Backtest Review

information_ratio -0.539 (negative active value vs benchmark); returns concentrated in a few outlier days (kurtosis 18.1)

Outcome Summary

After repeated single-name confluence ports failed the multiple-testing gates, this basket tried to fix the problem structurally by averaging the proven signal across BTC, ETH, and BNB on Binance's multi-year history, arguing that a diversified return stream would be harder to noise-select. The 6.6-year, 2,148-trade backtest was decisive but disproved the thesis on its own terms: the legs remained correlated, drawdowns compounded to 27% (CI high 83%), and the basket's 0.319 Sharpe undershot even the single ETH leg abandoned earlier in the session. With profit factor at 1.116, a negative information ratio, and a fatal $1.96M capacity ceiling from 22.58% impact cost, the analyst abandoned it at the backtest-review gate — no code bug to iterate and no edge that 3-parameter tuning could lift without curve-fitting.

Outcome Summary

Diversification only suppresses overfit if the legs are actually decorrelated — here three confluence legs on correlated majors stayed correlated (0.46-0.61), so the basket delivered a lower Sharpe than a single leg rather than a more stable objective; and a thin momentum edge that survives only at ~$2M capacity with a quarter of gross PnL lost to market impact is not promotable regardless of sample size.

Outcome Summary

The analyst abandoned it at the pre-optimization BACKTEST_REVIEW gate — optimization never ran — because the central diversification-as-robustness thesis was falsified by the backtest itself: the basket Sharpe 0.319 came in below the single-name ETH leg (0.43) already abandoned, since the legs stayed positively correlated (0.46-0.61) and drawdowns compounded. Profit factor 1.116 tripped the OHLCV-momentum PF<1.2 abandon signal, and a $1.96M capacity with 22.58% impact cost made the edge real only at toy scale.

Outcome Summary

A long-short, multi-instrument trend basket that ran the factory's proven 4H+1D momentum-confluence signal independently on three Binance USD-M majors (BTC, ETH, BNB) with three shared parameters and equal-risk legs, betting that averaging three decorrelated confluence streams would suppress overfit — lowering PBO and raising deflated Sharpe — where single-name variants kept failing the robustness gates.

Outcome Summary

Over a huge decisive sample — 2,148 two-sided trades (1,061 long / 1,087 short) across 6.6 years — it returned +170.9% but with a full-sample Sharpe of only 0.319 (CI [-0.29, 0.99] straddling zero), profit factor 1.116, information ratio -0.539, and 27.1% max drawdown. avg_trade_return_pct 0.459% cleared the fee floor, but impact_cost_pct was 22.58% with capacity of just $1.96M, and returns were concentrated in a few outlier days (kurtosis 18.1).
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.