CrossSectionalMomentumFactorNeutralLS
Hypotheses
Crypto Cross-Sectional Momentum Factor — Dollar/Beta-Neutral Long-Short: Rank the Liquid Perp Universe by Trailing 90-Day Return (Skipping the Last 7 Days), LONG the Top Quintile / SHORT the Bottom Quintile, Monthly Rebalance (BINANCE USD-M, Market-Neutral, 3-Parameter)
Hypotheses
A MARKET-NEUTRAL, MULTI-INSTRUMENT cross-sectional MOMENTUM factor — the single most-documented survivable anomaly in every asset class, and the CORRECT-SIGN, correct-horizon counterpart to the cross-sectional REVERSAL that was just exhaustively falsified here. It is distinct from every graveyard family: NOT single-name absolute trend/momentum (L62's 0.003-survival class — this is a beta-neutral relative sort, not a directional bet on one coin), NOT diversified TSMOM (abandoned overfit — that is absolute/always-in; this is a relative long-short that nets out market beta), NOT the falsified reversal (opposite sign, and it SKIPS the recent week where reversal lives), NOT calendar, NOT an options structure (which all die on Layer-3 timeouts), NOT a funding/OI z-score, NOT the always-in multi-factor composite of L60 (single factor). It is also distinct from my in-pipeline alt/BTC relative-strength basket: that trades four specific alt/BTC RATIOS with per-ratio trend signals and a shared BTC hedge, whereas this ranks the FULL ~20-name universe cross-sectionally by absolute trailing return and trades top-vs-bottom quintiles. MECHANISM: assets that outperformed their peers over the past 1-3 months continue to outperform over the next month (under-reaction / capital-chasing-performance), while the laggards keep lagging; a dollar- and beta-neutral long-top / short-bottom book monetizes that persistence with no market exposure. Standard '12-2'-style construction adapted to crypto: rank on the trailing 90-day return but SKIP the most recent 7 days to avoid contaminating the momentum signal with short-term reversal. DATA (named per L61, all present, TIMEOUT-SAFE): BINANCE USD-M 1D OHLCV for the universe only — trailing return and beta are O(1)-per-bar incremental statistics, NO full-series rescans, NO supplementary feed, NO options surface. Only 3 tunable parameters (momentum_lookback, quintile_fraction, rebalance_days).
Hypotheses
Implements the hypothesis exactly: skip-7 trailing-90d cross-sectional rank, top/bottom quintile, monthly calendar-anchored rebalance, dollar- AND beta-neutral, 3 tunables (momentum_lookback, quintile_fraction, rebalance_days), 1-DAY bars, no supplementary feed, no history rescan. HOWEVER I must report that the hypothesis is now falsified and I recommend ABANDON rather than a fifth iteration. QA asked for a trade-affecting change, so this iteration I tested the one construction family left untested -- risk-adjusted ranking (rank by momentum/vol) and inverse-vol leg weighting instead of equal notional -- engine-free on the implementable fixed 21-name universe of real BINANCE USD-M 1-DAY closes (2020-10-16..2026-08-06, 0.10% round trip charged on every leg of every rebalance), over mode(raw|risk) x volweight(0|1) x lookback(30/60/90/120) x hold(14/21/30) = 48 cells. Nothing clears noise: best cell +0.57%/period at t=+1.61 (risk-rank, inverse-vol, L=120/H=14), worst -1.02% at t=-1.47, 48-cell mean about -0.05%. I then settled the 'just ship L=120' question with a permutation null: re-running the identical 48-cell grid with the cross-sectional ranking RANDOMLY PERMUTED at each rebalance (100 draws) gives a median max|t| of 2.86 and a 95th percentile of 3.80, and p(max|t|_null >= 1.61) = 1.00 -- a random ranking beats the momentum sort's best cell 100% of the time on this grid. So every trade-affecting change available to me here is a mined cell that a coin flip outperforms, and I refuse to ship one; the defaults stay at the hypothesis's own spec (90/7/0.20/30). This is the fourth independent construction to come back flat or negative (raw-rank grid, quintile width, beta-hedge on/off, now risk-adjusted rank + inverse-vol weighting) and it agrees with the corrected-engine sandbox (-3.03% total, Sharpe -0.27, PF 0.73, avg_trade -2.30%). The three previously rejected rescues remain rejected: sign-flip is the just-falsified cross-sectional reversal, widening the universe reproduces a listing/survivorship artifact that the base class's alignment barrier makes unimplementable anyway, and grid-mining is now ruled out quantitatively by the permutation null. Code is therefore functionally unchanged from iteration 2 (docstring records the new evidence); venue stays BINANCE USD-M because the book is long+short, and leverage stays 1.0 because gross exposure is 0.60 of equity and nothing in sizing reads leverage.
Hypotheses
Negative-expectancy market-neutral cross-sectional momentum basket with a catastrophic drawdown, and the edge is proven indistinguishable from noise — not worth 2 hours of optimization. Over 243 trades: total_return -46.4%, CAGR -13.8%, avg_trade_return_pct -6.19% (NEGATIVE), profit_factor 0.79, expectancy -$230/trade, Sharpe -0.034 (CI [-0.80, 0.72] straddling zero), PSR 0.47, and max_drawdown 65.0% (CI to 93%) — past the L19 hard-abandon line — losing in 5 of 6 years (only 2022 positive). The developer's permutation null settles the mining question: re-running the identical 48-cell grid with the cross-sectional ranking randomly permuted gives a null median max|t| of 2.86, so a RANDOM ranking beats the momentum sort's best cell (t=1.61) 100% of the time — any apparent edge is selection bias, and four independent constructions plus the corrected sandbox all agree the factor is negative-to-noise. The result is also outlier/leverage-driven and mis-neutralized: return_kurtosis 86.9 with monthly swings of -67%/+126% (2021) and exposure_pct 879% despite the stated dollar-neutral 0.6x gross, with stressed-regime return -66%. This is the L45/L52 market-neutral cross-sectional momentum basket (0/84 survival); benchmark_meaningful is correctly false and the absolute risk-adjusted metrics are negative. Unlike a sandbox-only abandon, the full-history backtest CONFIRMS the developer's recommendation. No tuning of momentum_lookback/quintile_fraction/rebalance_days flips a negative-expectancy factor a permutation null shows is random. Failure pattern: no_edge/risk_reject market-neutral cross-sectional momentum (L45/L52).
Implementation
Dollar- and beta-neutral cross-sectional momentum factor on 21 liquid BINANCE USD-M perps (1-DAY bars). Each calendar-anchored rebalance slot it ranks every name by its trailing 90-day log return measured as of 7 days ago (the '12-2' skip that keeps short-term reversal out of the momentum signal), goes LONG the top 20% and SHORT the bottom 20% at equal USD notional per name (dollar-neutral by construction), and adds a BTC leg sized to offset the residual net beta of the two legs (rolling 90d covariance vs BTC, clamped 0.2-3.0), so the book is beta-neutral as well. Orders are netting deltas toward target notionals, so names that stay in the same quintile are not churned. Per-bar work is O(#names) bounded-deque updates; the ranking, betas and orders run only on the rebalance slot.
Verification Results
CLEAN RESTART 2026-09-04 — this run's verdict history and learning records were removed and it was restarted from verification. Its previous abandonment came from the pipeline, not from the market: the Layer-2 harness mis-bound @staticmethod helpers (fixed), QA issued terminal performance verdicts on an unoptimized smoke test (removed — QA now judges correctness only), and sandbox timeouts came from backtest-slot starvation (fixed). The hypothesis and the strategy code are unchanged. Verify the code on its merits; performance is decided later by the full backtest and the optimizer.
Verification Results
Route to the analyst for abandon-on-evidence at backtest-review; do NOT optimize (permutation null shows every cell is at or below random). Do not iterate further — the developer has closed the search across four constructions plus a permutation null, and the sign-flip is a separate (also dead) hypothesis for the Research Lead.
Verification Results
The edge is falsified, not the code. Cross-sectional momentum rank on this OHLCV universe is indistinguishable from a random ranking (developer's permutation null) and is a documented zero-survivor class (L81); the sandbox is a clear loser (-46.4%, Sharpe ~0, PF 0.84). The code faithfully and correctly implements the specified momentum factor, so this is a performance/edge determination for the analyst, not a defect, and it is not fixable within the hypothesis (the only profitable sibling is the sign-flipped cross-sectional reversal, itself just exhaustively falsified). The executable code is functionally unchanged from the prior iteration (docstring records the new permutation-null evidence), which is appropriate: no code change makes a false hypothesis true.
Backtest Review
Well-constructed, timeout-safe (243 trades, no supplementary feed); the correct-sign counterpart to the falsified reversal
Backtest Review
Rigorous falsification: developer's permutation null quantitatively rules out grid-mining
Backtest Review
Negative expectancy: total_return -46.4%, avg_trade_return_pct -6.19%, PF 0.79, Sharpe -0.034 (CI [-0.80,0.72]), expectancy -$230/trade
Backtest Review
max_drawdown 65.0% (CI to 93%) — past the L19 hard-abandon line; loses in 5 of 6 years
Backtest Review
Permutation null: a random ranking beats the momentum sort's best cell 100% of the time — edge is selection bias, not signal
Backtest Review
Outlier/leverage-driven (kurtosis 86.9, monthly -67%/+126%, exposure 879%) — neutralization did not hold; stressed regime -66%
Backtest Review
L45/L52 market-neutral cross-sectional momentum basket (0/84)
Iteration History
Verification failed (Layer 4 — QA review):
- The factor is not established on a survivorship-controlled universe, and both the developer's own study and the sandbox are negative (checklist #7). On the shipped fixed 19-name universe (every name present the whole 2020-10..2026-08 sample -- the only implementable set given the base class's alignment barrier), the hypothesis's own spec cell (momentum_lookback 90 / skip 7 / hold 30 / frac 0.20) returns -0.22% per period (t=-0.26, Sharpe -0.09), and across all 6 lookback x hold cells nothing exceeds |t|=1.5 (span -1.02% to +0.19%), with quintile width barely mattering. The corrected-engine sandbox confirms it: total_return -3.03%, Sharpe -0.27, profit_factor 0.73, avg_trade_return_pct -2.30%, over 25 trades. My corroboration bar is met on the negative side -- the honest study is flat-to-negative AND the sandbox is negative -- so there is no edge to advance. This is also the L7 cross-sectional-momentum-rank class (zero survivors), here confirmed dead on a clean crypto universe.
- The only positive result the developer found (+1.15% per period, Sharpe 0.61 on a wider 24-name universe admitting mid-sample listings SUI/ARB/APT/OP/INJ) is a listing/composition SURVIVORSHIP artifact, not the factor -- the gain comes from ranking names whose history begins inside the sample -- and it is also unimplementable here because the same-timeframe alignment barrier would defer every primary bar until the last-listed leg exists. The developer correctly refused to ship it. Flagged so the analyst does not resurrect that number as evidence.
Iteration History
Verification failed (Layer 4 — QA review):
- The factor is falsified and the developer concurs with abandonment; this is the second review and no trade-affecting change was made (checklist #7). Two independent measurements agree it has no edge: the developer's engine-free study on the implementable fixed 19-name universe puts the hypothesis's own spec cell at -0.22%/period (t=-0.26, Sharpe -0.09) with all six lookback x hold cells inside |t|=1.5, and the corrected-engine sandbox is independently negative (total_return -3.03%, Sharpe -0.27, profit_factor 0.73, avg_trade_return_pct -2.30%, 25 trades). My corroboration bar is met on the negative side -- both the honest study and the sandbox are negative. The developer correctly rejected all three rescues: sign-flip is the just-falsified cross-sectional reversal, widening the universe reproduces only a listing/survivorship artifact (unimplementable via the alignment barrier), and grid-mining the single +0.19% cell (t=+0.29 of six) is selection bias; dropping the beta hedge only moves -0.22% to -0.10% and breaks the titled beta-neutrality.
Iteration History
Verification failed (Layer 4 — QA review):
- No fee-clearing edge. The sandbox backtest is decisively negative: total_return -4.70%, Sharpe -0.40, profit_factor 0.69, win_rate 0.44, and avg_trade_return_pct -3.47% (checklist #7 requires per-trade return above the ~0.10% round-trip cost — here it is deeply NEGATIVE). The developer's own docstring/rationale concede the hypothesis is FALSIFIED across four independent constructions plus a corrected-engine study (-3.03% total, Sharpe -0.27, PF 0.73, avg_trade -2.30%) and a permutation null in which a RANDOM cross-sectional ranking beats the momentum sort's best cell 100% of the time (p(max|t|_null >= 1.61) = 1.00). The edge is absent, not merely weak.
- Statistically unmeasurable trade count (learning L16). Only 25 trades over 361 sandbox days from a monthly-rebalanced 21-name book — far below the ~100-trade floor needed to separate edge from noise. Sharpe CI -1.84..+1.37 straddles zero, probabilistic_sharpe 0.32, benchmark_meaningful=false. The point metrics are negative regardless.
Backtest and paper results are hypothetical. Trading involves risk of loss.