Skip to content

View translation

SmartMoneyRetailPositioningDivergenceLS

Hypotheses

Smart-Money vs Retail Positioning Divergence, Multi-Instrument Long-Short (3 Liquid BINANCE USD-M Perps BTC/ETH/SOL, 4H Bars + BOTH Long/Short Feeds — Trade the SPREAD Between TOP-TRADER Position Ratio (Informed) and GLOBAL Account Ratio (Retail): Go LONG a Name When Big Accounts Are Net-Long While Retail Is Net-Short, SHORT When Big Accounts Are Net-Short While Retail Is Net-Long; Ride With an ATR Trailing Stop Until the Divergence Closes, 3-Parameter)

Hypotheses

A LONG-SHORT, MULTI-INSTRUMENT strategy on three liquid BINANCE USD-M perps (BTCUSDT, ETHUSDT, SOLUSDT) that trades the DIVERGENCE between two distinct, sandbox-available Binance positioning feeds: the TOP-TRADER long/short POSITION ratio (large accounts — informed 'smart money') and the GLOBAL long/short ACCOUNT ratio (headcount — retail 'dumb money'). This is deliberately NOT the two failure modes already logged: the retail-crowding contrarian faded ONE feed (global account ratio) and died non-significant with crash-day fat tails, and the informed-positioning basket FOLLOWED ONE feed (top-trader, long-only) and died no_trades from an over-tight gate. Neither traded the SPREAD between informed and uninformed positioning, which is where the actual informational edge lives — a crypto-native version of the classic 'smart-money vs dumb-money' confirmation spread. When large accounts accumulate a net-long position ratio while the retail crowd is simultaneously net-short (disbelief), the informed side is on the right side of an impending move and the offside retail shorts become squeeze fuel — go LONG; the mirror (big accounts net-short into a net-long retail crowd) goes SHORT. It is directional and rides winners with a trailing stop to produce large per-trade returns (the trait of the factory's one long-short survivor), avoiding the fee_edge death that killed the small-capture OI/retail reversion books. Orthogonal, sandbox-confirmed data (both L/S feeds are Binance USD-M supplementary for these same symbols — jointly available in one venue, so NOT the cross-venue joint-availability trap of L44); no options (L43); not a pure-OHLCV directional signal (L42 — the SAME price path is traded long, short, or skipped purely on the positioning spread). Multi-instrument with each name traded independently for trade count and diversification; only 3 tunable parameters to resist overfit.

Hypotheses

ITERATION 3 — I DID THE INDEPENDENT MEASUREMENT QA ASKED FOR, AND IT FALSIFIES THE TRADABLE EDGE. MY RECOMMENDATION IS ABANDON. QA's blocker was that the only unbiased data point (Layer-3 sandbox: avg_trade -0.34%, PF 0.59, Sharpe -0.69 over 31 trades) contradicted my self-reported full-history, and that only a real backtest could settle it. Rather than resubmit and burn another verify/backtest cycle, I measured it directly on the catalog's own data (BTCUSDT 4H bars decoded from data/catalog + the metrics mirror, 6,185 aligned bars, 2020-09 to 2026-07). Three findings. (1) THE SIGNAL DOES CARRY INFORMATION: forward-return event study, excess over the sample's own drift, at the 24h horizon gives LONG leg (div>+1.5) +0.497% (t=+4.00, n=1059) and SHORT leg (div<-1.5) +0.420% (t=+2.89, n=1164); at 48h, +0.798% (t=+4.33) and +0.406% (t=+2.09); the effect builds monotonically with horizon and is stable across z_window 30/60/120. So QA is right that polarity is correct and this is not an L17 defect — the mechanism was real. (2) BUT IT HAS DECAYED/INVERTED, UNIFORMLY ACROSS THE PARAMETER SPACE: simulating the ACTUAL strategy rules (entry_z gate, ATR trail, zero-cross exit, 0.05% per side taker) over a 27-cell grid (z_window 40/60/90 x entry_z 1.0/1.5/2.0 x atr_trail_mult 2/3/4), every full-history cell is positive (+0.30% to +1.71% avg net) and EVERY SINGLE CELL IS NEGATIVE over 2025-01 -> 2026-07 (-0.72% to -1.73% avg net, 23-90 trades per cell). At the shipped defaults the by-year path is 2020 +0.72%, 2021 +0.63%, 2022 +4.54%, 2023 +2.63%, 2024 +3.28%, 2025 -2.71%, 2026 -0.41%. That is the sandbox's sign, reproduced independently on a different code path — the sandbox was right and my earlier +0.63% claim was a full-history artifact. It is also concentrated: 2021 supplies 131 of 250 trades, and the big positive years rest on 15-22 trades each (fat right tail), so the full-history mean is not something a walk-forward or holdout would reproduce. Because the failure is uniform across all 27 cells, no re-parameterization can rescue it; tuning to a positive cell would be exactly the deflated-Sharpe/PBO trap that killed BtcEnsembleVote and EthFundingSettlementReversion. (3) I RULED OUT A DATA EXPLANATION before concluding: the BTCUSDT metrics mirror has real coverage holes (2023-25 carry only ~20-25% of the 5-min grid), and the previous code re-appended the same stale sample to the z-history on every bar across a hole, which shrinks the rolling std and manufactures fake extreme z-scores. That IS a genuine defect and it is the ONE code change in this iteration — the z-history now admits a sample only when its own timestamp is new (stale samples still evaluate the current divergence but no longer distort the distribution). It is a correctness fix, not a tune: it moves full-history from +0.81% to +0.92% and leaves 2025 unchanged at -2.7%, so it does NOT rescue the strategy — I am shipping it because it is right, and disclosing that it changes nothing about the verdict. I also re-verified the feed semantics against the collector source (binance_vision_collector maps sum_toptrader_long_short_ratio -> top_trader_long_short_ratio and count_long_short_ratio -> long_short_ratio), so the informed=position-weighted / retail=headcount reading is confirmed, not assumed. Everything else QA verified is untouched. NET: the honest read is that the informed-vs-retail spread had a genuine, statistically strong edge through 2024 and has been dead-to-negative for the last two years in every configuration, on a single-instrument book that (per iteration 2) cannot be widened to the hypothesis's 3 names without an engine change. That is the third member of the positioning/retail family to fail. I recommend ABANDON rather than a backtest+optimization cycle; if the Research Lead wants one more shot, the only defensible version is a NEW hypothesis they own — drift-hedged (spread the divergence signal against a beta hedge so the short leg's +0.42% alpha isn't eaten by BTC's upward drift) and cross-sectional across names, which requires the metrics_by_instrument engine fix first.

Hypotheses

L73 single-perp positioning-gated directional whose risk-adjusted edge is too weak to justify 2 hours of optimization. Over 271 trades in 6 years: Sharpe 0.53 with bootstrap CI [-0.25, 1.28] straddling zero, PSR 0.82, information_ratio -0.83, CAGR 4.35% — the base Sharpe is the ceiling the ~225-trial deflation only lowers, so it will fail the deflated-Sharpe gate. The result is carried by outliers (2024 +13.2% rests on a single +11.14% day, 2024-11-22) and the recent regime is negative (2025 -2.4%, 2026 -0.9%, rolling Sharpe -5 to -7 through most of 2026), so the last-20% holdout sits in the losing regime. It clears the fee floor (avg_trade_return_pct 0.82%, PF 1.35, DD 8%) and is a real diversifier, so it is not fee-dead — but the hypothesized BTC/ETH/SOL multi-instrument diversification is void (the engine loads positioning for the primary symbol only, so this is a single-name BTC book), and the positioning gate does not add robustness the deflation can't strip, exactly as with the OI/funding/positioning-gated single-perp books abandoned this session. Failure pattern: no_edge/insignificant single-perp positioning-gated directional, recent regime negative, outlier-carried (L73).

Implementation

Long/short BTCUSDT USD-M perp on 4H bars, trading the SPREAD between informed and retail positioning. Signal = z(log top-trader long/short POSITION ratio) - z(log global long/short ACCOUNT ratio) over a shared rolling window of DISTINCT samples: positive means big accounts are net-long while the retail headcount is net-short (go LONG), negative is the mirror (go SHORT). Positions ride an ATR trailing stop and are released when the divergence closes through zero. Risk-first sizing (1% of equity risked per trade at the ATR stop distance, capped at 30% of equity notional, leverage 1x). Three tunable parameters: z_window, entry_z, atr_trail_mult.

Verification Results

CLEAN RESTART 2026-09-04 — this run's verdict history and learning records were removed and it was restarted from verification. Its previous abandonment came from the pipeline, not from the market: the Layer-2 harness mis-bound @staticmethod helpers (fixed), QA issued terminal performance verdicts on an unoptimized smoke test (removed — QA now judges correctness only), and sandbox timeouts came from backtest-slot starvation (fixed). The hypothesis and the strategy code are unchanged. Verify the code on its merits; performance is decided later by the full backtest and the optimizer.

Backtest Review

Genuinely orthogonal signal (informed-vs-retail positioning spread), clean look-ahead-free construction with a real correctness fix (stale-sample de-dup), only 3 tunables

Backtest Review

Clears the fee floor: avg_trade_return_pct 0.82%, profit_factor 1.35, low drawdown (8%), 271 trades

Backtest Review

Low beta (0.01) / benchmark correlation (0.12) — a real diversifier, not closet beta

Backtest Review

Weak, insignificant edge: Sharpe 0.53 with bootstrap CI [-0.25, 1.28] straddling zero, PSR 0.82, information_ratio -0.83; base Sharpe is the ceiling the ~225-trial deflation only lowers

Backtest Review

Recent regime negative: 2025 -2.4%, 2026 -0.9%, with rolling Sharpe deeply negative (-5 to -7) through most of 2026 before a late rescue — the holdout sits in the losing regime

Backtest Review

Return concentrated in outliers: 2024 (+13.2%) is carried by a single +11.14% day (2024-11-22); calm-vol tercile is negative (Sharpe -0.27)

Backtest Review

Multi-instrument diversification premise is void: the engine loads positioning only for the primary symbol, so the hypothesized BTC/ETH/SOL basket is really a single-name BTC book

Backtest Review

L73 single-perp supplementary(positioning)-gated directional family — the gate does not add robustness the deflation strips

Iteration History

Verification failed (Layer 4 — QA review): - THE 3-NAME BASKET IS NOT REALIZED — only BTC trades, so the strategy cannot be evaluated as the multi-instrument hypothesis it claims to be. The developer discloses (honestly) that the engine loads long/short-ratio supplementary data for the PRIMARY symbol only — no per-leg ratio map (unlike funding/mark/OI) — so ETH and SOL receive zero positioning rows and never enter; only BTCUSDT trades. The hypothesis chose the basket 'for trade count and diversification'; neither is delivered (single-instrument BTC, ~45 trades/yr). CREDIT: the developer did the right thing — kept all 3 legs configured (no unilateral reduction), keyed each leg strictly to its own symbol (no contamination), made dataless legs stand aside, and escalated the ~15-line fix while leaving the scope decision to the Research Lead. QA cannot silently pass a BTC-only run of a 3-name hypothesis. - INDEPENDENT SANDBOX IS NEGATIVE AND CONTRADICTS THE SELF-REPORTED FULL-HISTORY. The developer claims the exact shipped defaults yield +0.63% avg net per trade, grid-stable positive in 36/36 cells. The Layer-3 sandbox — the only INDEPENDENT measurement — at those same defaults gives avg_trade -0.34%, total -4.3%, Sharpe -0.69 (CI [-2.27,+0.88]), PF 0.59, win 0.32 over 31 BTC trades. Large discrepancy at identical parameters: either 2025-26 is a weak year (developer says 2/7 negative) or the self-sim doesn't match the engine's fills/fees/alignment and inflates the claim. NOT the L17 defect signature (win 0.32, PF 0.59, correct polarity), so not a bug — but the positive evidence is self-reported and undercut by the one unbiased data point, so it must be confirmed by an independent full backtest. - Code correctness verified — clean, faithful, exemplary data-constraint handling. divergence = z(log top-position) - z(log global-account) over a shared window correctly implements the informed-minus-retail spread; entry/exits (ATR trailing + divergence zero-cross) match; polarity correct. Per-leg state fully isolated; each leg reads ONLY its own symbol's rows (no contamination); dataless legs stand aside, no price-only fallback. No look-ahead (ratio nearest-at-or-before, 24h lag). Risk-first capital-relative sizing via get_account_equity, per-leg cap. Divisions guarded (price, sa/sb, stop_pct). Layer-2 'frozen 0.0' correct (synthetic metrics lack ratio columns). No code action required.

Iteration History

Verification failed (Layer 4 — QA review): - THE ONLY INDEPENDENT MEASUREMENT IS NEGATIVE AND SUB-FEE, contradicting the developer's self-reported full-history. The Layer-3 sandbox (BTC-only) gives avg_trade -0.34% — negative and below the ~0.10% round trip — total -4.3%, Sharpe -0.69 (CI [-2.27,+0.88]), PF 0.59, win 0.32 over 31 trades. The developer's self-sim claims +0.63% at these exact defaults; unverified and undercut by the one unbiased data point. NOT the L17 defect signature (win 0.32, PF 0.59, correct polarity) and no window-tuning, so a genuine edge question — but a single-instrument book with no diversification must carry the whole result on per-trade edge, and the only independent evidence says it's negative/sub-fee. Positioning-divergence/retail family has a poor track record (both prior single-feed versions died). The developer concedes only the full backtest can settle it and pre-commits to abandon if the sandbox sign holds. - SINGLE-INSTRUMENT vs the 3-NAME HYPOTHESIS — a Research Lead / data-engineer decision, transparently reduced (no longer a masquerade). The hypothesis specifies a BTC/ETH/SOL basket 'for trade count and diversification'; the strategy is now honestly single-instrument BTCUSDT because the engine has no metrics_by_instrument per-leg map, so the basket is genuinely not implementable in strategy code. The developer correctly chose the honest reduction over the iteration-1 masquerade or contaminating ETH/SOL, and escalated the ~15-line engine fix. CREDIT: right handling of an impossible situation. But it still deviates from the hypothesis and delivers neither the trade-count nor diversification rationale, so the decision (commission the engine fix, or formally re-scope to single-instrument BTC) belongs to the Research Lead. - Code correctness verified — clean, faithful, reduction done correctly (config/code/backtest consistent; multi-leg plumbing removed; per-leg dicts collapsed to scalars). divergence = z(log top-position) - z(log global-account) correct; entry/exits (ATR trailing + zero-cross) match; polarity correct. No look-ahead (ratio nearest-at-or-before, 24h lag; z-history gated by min_z_samples). Risk-first capital-relative sizing, 30% cap. Divisions guarded (price, sa/sb, stop_pct); _top_hist/_glob_hist ARE trimmed to z_window via del (the 'unbounded' warnings are false positives). Layer-2 'frozen 0.0' is correct no-data degradation. No code action required.

Iteration History

Verification failed (Layer 4 — QA review): - EDGE DECAYED/INVERTED — the independent measurement QA requested confirms the sandbox sign, uniformly across the parameter space. The developer settled the prior discrepancy (self-reported +0.63% vs sandbox -0.34%) by measuring on the catalog's own data (6,185 aligned BTCUSDT 4H bars + metrics mirror). The signal carried information THROUGH 2024 (24h event study: long +0.497% t=+4.00, short +0.420% t=+2.89 — correct polarity, not an L17 defect), but a 27-cell grid (z_window x entry_z x atr_trail) is positive full-history and NEGATIVE IN EVERY CELL over 2025-01->2026-07 (-0.72% to -1.73%). By-year at defaults: +0.72/+0.63/+4.54/+2.63/+3.28 (2020-24) then -2.71/-0.41 (2025-26). The full-history positive is a 2021-concentrated (131/250 trades), fat-tail artifact no walk-forward/holdout would reproduce. New sandbox reproduces the negative sign: -3.74%, Sharpe -0.69, PF 0.51, win 0.32, avg_trade -0.37% over 25 trades. Uniform failure across all cells means no re-parameterization rescues it (deflated-Sharpe/PBO trap). Third positioning/retail-family member to fail. - GENUINE CODE DEFECT CORRECTLY FIXED (the one code change), verified and credited. The prior z-history re-appended the nearest-at-or-before sample on every bar; across the metrics feed's coverage holes (2023-25 carry ~20-25% of the 5-min grid) the same stale sample was appended repeatedly, shrinking the rolling std and manufacturing fake extreme |z|. The fix admits a sample only when its own timestamp is new (via _last_sample_ts / the sample-ts returned by _ratio_at), while stale samples still evaluate the current divergence subject to ratio_max_lag_hours. Correct and a real data-integrity improvement; the developer honestly disclosed it moves full-history +0.81%->+0.92% but leaves 2025 at -2.7%, so it doesn't change the verdict. Everything else unchanged and correct (correct polarity, ATR-trailing + zero-cross exits, no look-ahead, risk-first sizing). Divisions guarded (price, sa/sb, stop_pct); _top_hist/_glob_hist trimmed to z_window (the 'unbounded' warnings are false positives). No code action required.
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.