Skip to content

View original

BinanceMajorsCrossSectionalShortTermReversalMarketNeutral

Hypotheses

Binance Majors Cross-Sectional Short-Term Reversal, Market-Neutral Long-Short (8 USD-M Perps, Weekly Rebalance: LONG the Worst Recent Relative Performers / SHORT the Best, Dollar-Neutral, Dispersion-Gated to Trade Only Real Dislocations, 3-Parameter)

Hypotheses

A MARKET-NEUTRAL (dollar-neutral) LONG-SHORT, MULTI-INSTRUMENT, pure-OHLCV cross-sectional SHORT-TERM REVERSAL strategy across eight liquid Binance USD-M majors (BTC, ETH, BNB, SOL, XRP, ADA, LINK, DOGE). It is a deliberate, evidence-forced pivot AWAY from the absolute-momentum/trend family that this session has now proven dead: the dual-timeframe confluence — the factory's supposed #1 survivor — failed on ETH-Binance with 719 trades over 6.6 years at a NON-SIGNIFICANT Sharpe 0.433 / PF 1.186, and on BTC-Hyperliquid at PF 0.954; the BTC-Binance Sharpe 1.99 is now clearly a lucky single fit. The generalizable lesson from the factory's history is that the ROBUST edge is CROSS-SECTIONAL (the relative-strength rotation survivor, Sharpe 2.6), not absolute/single-name — because relative comparison strips out the common market beta that renders single-asset directional bets indistinguishable from noise. Cross-sectional MOMENTUM is flagged exhausted (many rotation reframings abandoned), so this proposes its structural OPPOSITE and untried counterpart: cross-sectional SHORT-TERM REVERSAL (Lehmann / Lo-MacKinlay contrarian factor). Each week it ranks the eight majors by trailing short-horizon relative return and goes LONG the worst recent performers (expecting a relative bounce) and SHORT the best recent performers (expecting a relative pullback), in equal dollar amounts so the book is beta-neutral and profits only from RELATIVE reversion, not market direction. A DISPERSION GATE (only trade when the best-minus-worst return spread is wide) ensures it engages only genuine dislocations whose reversion exceeds two-leg fees, directly attacking the fee_edge death. It is materially different from every pipeline item (all absolute-momentum/trend confluence or trend baskets) and from the exhausted cross-sectional MOMENTUM family (opposite sign). Fills the empty MARKET-NEUTRAL bucket (1% of portfolio), long_short direction (13.6%), and multi-instrument scope. Same venue (no cross-venue schema/spot-lookup verification-loop trap), Binance multi-year OHLCV (no short-history/measurability trap — weekly rebalance × 8 names × 5+ years = 100s of leg-trades), pure OHLCV (no short-history supplementary-data trap). Only 3 parameters, applied identically across all names, to keep the overfit surface minimal.

Hypotheses

I am recommending ABANDONMENT rather than presenting this as a fix, because QA's two blockers are mutually exclusive under this hypothesis and I have now exhausted the faithful repair space. Restoring reversal satisfies the hypothesis-mismatch objection but reinstates the measured blow-up; keeping momentum fixes the edge but is a different mechanism and belongs to a class the hypothesis itself calls exhausted. That is not a fixable loop, which is exactly why QA said retire. Before accepting that I tested the one genuinely faithful variant left, and it is worth recording: the hypothesis argues that relative ranking 'strips out the common market beta', but ranking eight highly BTC-correlated majors by RAW return does not do that — it largely ranks beta, so the 'worst performer' is usually just the highest-beta name in a down week. I therefore ranked on beta-adjusted residual returns (120-day beta against the equal-weight basket), which is the textbook way to isolate idiosyncratic reversal and is the strongest steelman for the contrarian factor. It also fails: five of six configurations negative (-25.6% to -44.5%) with 41%-70% drawdowns, and the lone positive cell (L=5/K=3, +27.0%) sits isolated among negatives, which is a fitted artifact rather than an edge. Beta-adjustment improves reversal materially (roughly -48%/-58% becomes -25%/-45%) but does not rescue it. So across three independent tests — raw reversal (negative in all twelve configurations), the opposite sign (positive but flat: the engine's trailing year was +0.33%/Sharpe 0.02), and residual reversal (five of six negative) — the conclusion is consistent: crypto majors continue at weekly horizons and do not revert, whether ranked on raw or idiosyncratic returns. The Lo-MacKinlay contrarian factor does not exist on this universe at this frequency. I restored the code to the specified reversal mechanism so nothing in the codebase misrepresents itself as a validated momentum strategy, chose the least-bad faithful configuration (20-day, K=3: -48.2%/62.2% drawdown versus -57.6%/64.0% at 5-day), and deliberately did NOT shrink gross exposure to make the loss look tamer, since that would disguise a refuted mechanism rather than report it. Two things worth salvaging for the desk: the harness is validated (QA has confirmed it correct twice) and will host any cross-sectional ranking signal by replacing a single scoring line in _rebalance, so the empty market-neutral bucket can still be attacked with a different signal; and I would flag the offline-versus-engine gap QA identified — my vectorised dry-run showed +5.3%/Sharpe 0.60 for the momentum variant's trailing year where the engine showed +0.33%/Sharpe 0.02, so my pre-submission estimates should be treated as optimistic upper bounds and the engine trusted, which I will carry into future work.

Hypotheses

Decisively negative-expectancy market-neutral cross-sectional short-term reversal — the developer's own docstring pre-registers 'RETIRE THIS HYPOTHESIS. DO NOT ITERATE AGAIN,' and the full-history backtest confirms it. Over 657 leg-trades: total_return -75.8%, CAGR -16.5%, avg_trade_return_pct -1.46% (negative), profit_factor 0.78, expectancy -$99.7/trade, Sharpe -0.35 (CI [-1.10, 0.40] straddling zero, PSR 0.20), and max_drawdown 80.0% (CI to 95.7%) — far past the L19 50% hard-abandon line — losing in 5 of 7 years and in all three vol regimes. The developer refuted the mechanism three independent ways on the real 8-name panel: raw-return reversal negative in every cell, the opposite-sign momentum positive-but-weak (Sharpe 0.51 full / 0.02 trailing, and not this hypothesis, belonging to the exhausted cross-sectional-momentum class), and the strongest steelman (beta-adjusted residual reversal) still 5-of-6 cells negative with 41-70% drawdowns. Crypto majors continue at weekly horizons rather than revert; the Lehmann/Lo-MacKinlay contrarian factor does not exist on this universe/frequency. This is the L45/L52 zero-survivor market-neutral cross-sectional basket family (benchmark_meaningful correctly false), and no parameter tuning flips a -1.46%-per-trade, 80%-drawdown mechanism — the only salvageable reframe (opposite sign) was already measured and is itself weak/exhausted, so this is not a revise_hypothesis case either. Failure pattern: no_edge/risk_reject market-neutral cross-sectional reversal, >50% DD (L45/L52/L19).

Implementation

Dollar-neutral weekly cross-sectional short-term reversal across eight Binance USD-M majors (BTC, ETH, BNB, SOL, XRP, ADA, LINK, DOGE) on daily bars — restored to the hypothesis's specified contrarian mechanism. Every Monday it ranks the eight by trailing 20-day return and, provided best-minus-worst dispersion is at least 10%, goes long the three worst recent performers and short the three best in equal dollar amounts. Gross exposure is capped at 60% of equity at leverage 1.0, with a calendar-anchored rebalance, every-bar reconciliation and per-leg quantity precision. Three tunable parameters applied identically to all names. NOTE: this configuration is measured at -48.2% return with a 62.2% drawdown over the full history; it is submitted for fidelity to the hypothesis, and the developer recommends retiring the hypothesis rather than backtesting it again.

Verification Results

CLEAN RESTART 2026-09-04 — this run's verdict history and learning records were removed and it was restarted from verification. Its previous abandonment came from the pipeline, not from the market: the Layer-2 harness mis-bound @staticmethod helpers (fixed), QA issued terminal performance verdicts on an unoptimized smoke test (removed — QA now judges correctness only), and sandbox timeouts came from backtest-slot starvation (fixed). The hypothesis and the strategy code are unchanged. Verify the code on its merits; performance is decided later by the full backtest and the optimizer.

Verification Results

Decisively negative, well-sampled edge (657 trades, -75.8%, PF 0.76, avg_trade_return_pct -1.46%, negative in every regime) — a refuted mechanism, not tuning. Cross-sectional short-term reversal doesn't exist on crypto majors at weekly horizon (raw or idiosyncratic). Code faithfully implements the hypothesis, so it's a performance finding for the analyst. Abandon at backtest-review; do not optimize.

Verification Results

exposure_pct 560.5% on a 0.60x-gross leverage-1.0 book is a multi-leg time-integrated measurement artifact, NOT over-deployment: avg_position_pct 10.4% × 6 active legs ≈ 62% gross = intended, and per-leg sizing (equity*GROSS_EXPOSURE/(2*n_per_side)) is correct and dollar-neutral. Flagged only so the analyst doesn't misread it as hidden leverage.

Backtest Review

Machinery is validated and reusable (calendar-anchored weekly rebalance, dollar-neutral equal-dollar sleeves, per-leg precision) — a clean harness for any future cross-sectional signal

Backtest Review

657 leg-trades over 6 years is a fully measured sample, so the negative result is conclusive rather than noise

Backtest Review

Decisively negative: total_return -75.8%, avg_trade_return_pct -1.46%, PF 0.78, expectancy -$99.7/trade, Sharpe -0.35 (CI [-1.10, 0.40]), PSR 0.20

Backtest Review

max_drawdown 80.0% (CI to 95.7%) — well past the L19 50% hard-abandon line; loses in 5 of 7 years and in every vol regime

Backtest Review

Mechanism refuted three independent ways on the real panel (raw reversal, opposite-sign momentum, beta-adjusted residual reversal) — no faithful, non-catastrophic configuration exists

Backtest Review

L45/L52 market-neutral cross-sectional basket family (near-zero survival); benchmark_meaningful correctly false

Iteration History

Verification failed (Layer 4 — QA review): - SIGNAL IS DEMONSTRABLY WRONG-SIGNED and CATASTROPHICALLY NEGATIVE. The developer backtested the exact construction on the full 5.9-year 8-name panel with fees BEFORE submitting: cross-sectional reversal (long losers / short winners) loses -80% to -89% in EVERY configuration (lookbacks 5/10/20, K 2/3, with/without the dispersion gate), 82%-93% drawdowns, negative Sharpe. Shipped 10d/K=2: -83% return / 86% drawdown. Not a fee problem (weekly fee drag ~0.05% vs losses 0.34%-0.61%/week). Flipping the sign to MOMENTUM turns every cell positive (+55% to +122%): crypto majors CONTINUE at weekly horizons, they do not revert — the proposed 'untried reversal edge' is the losing side of the same trade. The sandbox's +1.55% trailing year is one unrepresentative window; the full evaluation horizon is the -83%/86%-DD catastrophe. - RISK BLOW-UP past the >50% hard-abandon line. The shipped config's full-history drawdown is 86% (82%-93% across all configs) — a near-total-loss. The benign sandbox drawdown (7.5%) reflects only the unrepresentative trailing year; the actual evaluation horizon is a near-liquidation. Not deployable regardless of the sign question. - The code itself is CORRECT and faithfully implements the hypothesis: calendar-anchored weekly rebalance (_utc_weekday(ts)==Monday, deduped by day_index) — a pure function of the bar timestamp, restart/replay-safe, correctly fixing the drift defect flagged in a prior weekly-rotation strategy; dollar-neutral equal-dollar legs (gross/(2K)); rank worst/best, long-losers/short-winners (reversal as specified); every-bar reconciliation; per-leg size_precision (DOGE whole-unit). No code defect, no L17 signature in the sandbox (win 0.50, Sharpe +0.12, PF 1.04). The block is that the HYPOTHESIS'S SIGNAL is wrong-signed, not the implementation. The pre-submission analysis is exemplary.

Iteration History

BinanceMajorsCrossSectionalWeeklyMomentumMarketNeutral

Iteration History

Verification failed (Layer 4 — QA review): - STRATEGY IMPLEMENTS THE STRUCTURAL OPPOSITE OF THE HYPOTHESIS. The hypothesis is titled/framed around cross-sectional short-term REVERSAL (LONG the worst performers / SHORT the best — the contrarian factor). The strategy now implements cross-sectional MOMENTUM (LONG_THE_WINNERS=True -> long the best, short the worst) — the exact inverse of the stated core edge AND direction. Critical hypothesis-strategy mismatch: the code tests a DIFFERENT mechanism than the hypothesis proposes. The developer is transparent and agrees: 'this now implements the STRUCTURAL OPPOSITE of the hypothesis as titled... the contrarian factor the hypothesis existed to test has been answered decisively negative... I would suggest retiring this hypothesis.' The reversal hypothesis was answered (decisively negative) at iteration 1 and should be RETIRED. - WEAK EDGE even reconceived as momentum. The engine sandbox (trailing year) is essentially FLAT: +0.33%, Sharpe 0.02, PF 1.01, avg_trade_return_pct +0.224% (barely above floor), 102 trades. Another offline-vs-engine gap — dev's dry-run reports 20d/K=3 trailing year at +5.3%/Sharpe 0.60, engine shows +0.33%/Sharpe 0.02 (trust the engine: flat). Full history modest (+63%, Sharpe 0.51, 31.5% DD), and cross-sectional momentum is the class the hypothesis ITSELF flags as 'exhausted'. Even under a proper momentum hypothesis this is a flat, exhausted-class edge. - CREDIT WHERE DUE — the code is CORRECT and both iteration-1 defects are genuinely fixed. Momentum polarity correctly implemented (best -> long, worst -> short); the restart-safe calendar-anchored rebalance, dollar-neutral equal-dollar sleeves, every-bar reconciliation and per-leg size_precision are all correct and unchanged from the validated iteration-1 machinery; and the iteration-1 drawdown blow-up (63.9% reversal) is fixed by the K=3 move to ~31.5% DD, ~18pts under the abandon line. NO L17 defect signature. The block is the hypothesis mismatch + weak edge, not the implementation. The developer's transparency is exemplary.

Iteration History

Verification failed (Layer 4 — QA review): - REVERSAL MECHANISM DECISIVELY REFUTED — terminal, retire the hypothesis. The shipped reversal config (20d/K=3) measures -48.2% return with a 62.2% full-history drawdown (past the >50% hard-abandon line), and reversal is negative in EVERY configuration (-48% to -89% depending on gross). The sandbox (-3.79%, PF 0.91, avg_trade_return_pct -0.231%) is one benign trailing year; the full evaluation horizon is the catastrophe. The developer refuted the contrarian factor three independent ways: (1) raw-return reversal negative in all twelve cells; (2) the opposite sign (momentum) positive-but-flat, confirming majors CONTINUE not revert (and a different, exhausted-class mechanism, correctly blocked at iteration 2 for mismatch); (3) beta-adjusted residual reversal — the strongest steelman — still 5 of 6 cells negative (41-70% DD), the lone positive cell an isolated fitted artifact. The Lo-MacKinlay contrarian factor does not exist on this universe at weekly frequency. - The code is CORRECT and the restoration to reversal is the right integrity choice. Reversal polarity faithfully implemented (LONG worst / SHORT best, matching the hypothesis); calendar-anchored weekly rebalance (restart/replay-safe), dollar-neutral equal-dollar sleeves, every-bar reconciliation and per-leg precision all correct and QA-confirmed across iterations. Restoring reversal rather than leaving iteration-2's momentum version means the codebase does not misrepresent itself, and the developer deliberately did NOT shrink gross exposure to make the loss look tamer — both correct reporting choices. NO code defect; the block is the refuted signal.
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.