Skip to content

View original

BtcTailRegimeConvexDirectionalLS

Hypotheses

BTC Binance USD-M Tail-Regime Convex Directional — Long-Short, Engage ONLY on Confirmed Large-Magnitude Directional Regimes, Flat in All Normal/Quiet Conditions, Wide Chandelier Ride (Daily Bars, 2-Parameter)

Hypotheses

A LONG-SHORT, single-instrument, pure-price CONVEX directional strategy on BTCUSDT.BINANCE (USD-M perpetual), daily bars. It is deliberately the OPPOSITE of the two directional templates that died this session: it is NOT always-invested on the trend sign (BTC macro-daily TSMOM died Sharpe 0.55, flat-CI), and it is NOT a goldilocks-vol gate that trades in MODERATE vol (that died verification_loop). Instead it is FLAT the large majority of the time and only ENGAGES when the market enters a confirmed LARGE-MAGNITUDE directional regime — a trailing return whose magnitude exceeds a high threshold AND accompanied by expanding volatility. This targets the one structural feature crypto reliably exhibits: fat-tailed, self-reinforcing TRENDING blow-offs (parabolic bull legs) and cascading crashes, where directional persistence is strongest and largest. By trading ONLY these rare, high-conviction legs and riding them with a wide chandelier, per-trade capture is large (fees trivial) and the book is decorrelated from the always-on trend factor (flat most days). This directly answers the session's core findings: (1) continuous trend on any single name lands in the Sharpe 0.3–1.3 dead-zone because most bars are noise/chop that dilute the edge and bleed fees — engaging ONLY on the extreme legs removes that dilution; (2) the deflated-Sharpe bar is a DOF function — this uses only 2 parameters (magnitude threshold, chandelier multiple), giving a low best-of-N hurdle. NOT a mean-reversion fade (L53 — it trades WITH the extreme move), NOT a squeeze/compression breakout (L54 — it engages on LARGE already-confirmed magnitude moves, documented to run, not low-vol pops), NOT a cross-sectional basket (L52), NOT a NON-PRICE-feed gate (magnitude + realized-vol are 100% price — NOT L46), NOT options/COIN-M (L50/L51), NOT a long-only basket (L48 — long-SHORT). Runs on BTC (only asset with demonstrated tradeable structure; cleanest fat-tail regime behavior). Risk profile: vol-targeted sizing on engagement (~2% equity risk anchored to the chandelier distance), capped at 25% of equity*leverage; 2x leverage cap (reads self.config.leverage). Symmetric engagement (both bull blow-offs and bear crashes) ensures trades fire across the full sample, not only in one regime.

Hypotheses

Iteration 3 fixes BOTH Layer-4 objections with one change: the bar type returns to 1-DAY (the hypothesis of record), and the tail window is shortened from 20 days to 5 so the gate still selects genuinely EXTREME moves while producing more distinct engagements. A |z| >= 1.25 move compressed into five days with expanding volatility is a blow-off or a cascade -- arguably more tail-like than the same z spread over a month -- and per-trade capture confirms it: +2.47% net per trade, versus +0.85% for the 4H version QA correctly called 'ordinary momentum' and +4.15% for the original 20-day version. Time-in-market is 25%, so the book is still flat three quarters of the time. Only two constants changed from the code QA already verified (lookback default 20 -> 5, entry_z default 1.0 -> 1.25 with the floor restored to 1.0); mechanism, guards, exits, sizing and both tunables are untouched. Layers 1-3 re-verified locally: static clean, all six synthetic scenarios pass, sandbox places 10 trades from 5 entries with no drops. RESULTS (replaying THIS EXACT CLASS over 2410 daily bars, 0.10% round-trip fees): 74 round trips (~11.2/yr) versus 27 (~4.1/yr) in iteration 1 -- 2.7x more observations while staying daily and staying extreme -- avg +2.47% per trade, profit factor 2.14, win rate 45%, kurtosis 2.0, top-3 concentration 27.6% (down from 60%), and positive in six of seven calendar years (2020 +4.90%, 2021 +5.91%, 2022 +0.34%, 2023 +3.66%, 2024 +0.26%, 2025 -0.40%, 2026 YTD +5.61%), with the trailing 365 days at +1.26%/trade. The real sandbox year returns +2.65% with Sharpe 0.94, PF 4.23, max drawdown 1.0%. TWO THINGS THE ANALYST MUST WEIGH, stated plainly because three iterations have now mapped the frontier exactly. (1) The measurability gap is NARROWED, NOT CLOSED: 74 trades clears the ~30-trade unmeasurable line but sits below the ~100-trade floor, and at 11 trades/yr a 15-day holdout still expects only ~0.46 trades, so it may be empty and trip the zero-trade gate. This is structural, not a coding choice -- the frontier is: daily + 20-day window = 27 trades (rejected as unmeasurable); daily + 5-day window = 74 trades (this iteration, still extreme, per-trade +2.47%); 4H = 235 trades but ordinary 1-sigma moves and +0.85%/trade (rejected as edge mismatch). A single asset on daily bars cannot produce ~100 rare-regime engagements; only pooling the identical rule across several majors could, which the hypothesis's single-instrument scope forbids. (2) The chandelier tunable is now effectively INERT -- 4.0 and 8.0 give identical results (74 trades, +2.47%) because with a five-day window the z-sign-flip exit almost always fires before the 6x-ATR give-back. That makes this a de facto ONE-parameter strategy: good for the deflated-Sharpe hurdle, but it means optimization has only entry_z to work with (1.0 -> 101 trades/+1.09%, 1.25 -> 74/+2.47%, 1.5 -> 56/+3.21%, 2.0 -> 37/+3.99%, a clean monotone frequency-vs-capture trade-off). My recommendation: if the holdout gate can be evaluated on pooled full-sample statistics rather than a 15-day window, this configuration is the best available reconciliation of the hypothesis's edge and the factory's measurement machinery; if not, the Research Lead should either amend the hypothesis to a multi-asset pooled version or retire it, because no single-asset daily parameterisation reaches 100 trades while remaining a tail-regime strategy.

Hypotheses

The convex tail-regime idea is genuinely novel and decorrelated (beta 0.009) with a real per-trade edge (PF 2.03, avg_trade_return_pct 2.23% well above fees), making it a cleaner result than the losing calendar sleeves — but it fires far too rarely to be validated. Only 72 trades over 6.6 years (~11/yr), well below the 100-trade measurability floor, and with return_kurtosis 26.3 the Sharpe 0.675 is outlier-driven (a handful of days — 2026-02-10 +2.73%, 2023-10-30 +2.45%, 2024-11-25 +2.07% — carry it); sharpe_ci_low 0.0516 sits right at zero, not distinguishable from no-skill. The rare cadence breaks the validation framework structurally: at ~11 trades/yr the 15-day holdout expects ~0.5 trades and is very likely EMPTY (a hard zero-trades gate), while the walk-forward OOS windows would be sparse and noise-dominated. Optimizing the 2 params over ~72 fat-tail events is best-of-N fitting to a few blow-offs, and post-optimization deflated Sharpe would very likely fail from a base whose CI touches zero. The economic edge is also tiny (+14.2% over 6.6y, ~2%/yr, information_ratio -0.71 vs buy-hold BTC, 2025 negative). Optimization cannot fix a cadence/measurability mismatch. Consistent with abandoning the low-trade-count rare-event strategies. Abandon at BACKTEST_REVIEW rather than spend the optimization budget.

Implementation

Long-short convex tail-regime strategy on BTCUSDT.BINANCE USD-M perpetual, DAILY bars, flat unless an extreme burst is confirmed. calculate_signal returns z = (log C[t] - log C[t-5]) / (daily sigma * sqrt(5)) every bar -- the trailing FIVE-day move in daily sigmas, continuous and varying. The book ENGAGES only when |z| >= 1.25 AND volatility is EXPANDING (20-day stdev above the 60-day baseline), i.e. a >1.25-sigma move compressed into a single week with rising vol -- a parabolic burst or a cascade -- and then takes the side of z; otherwise FLAT (time-in-market 25%). EXIT on a wide 6x-ATR chandelier from the best close since engagement or when z flips sign. Sizing risks ~2% of equity at the chandelier distance, capped at 25% of equity x 2x leverage. Two tunables: entry_z and chandelier_atr.

Verification Results

Verification failed (Layer 4 — QA review): - Hypothesis-vs-implementation mismatch on BOTH the stated timeframe and the core edge. (1) Timeframe (checklist item 1): the hypothesis of record is explicitly 'Daily Bars' (title and description repeatedly say daily), but the config bar_type is now BTCUSDT.BINANCE-4-HOUR — a direct hypothesis-vs-config timeframe mismatch. (2) Core edge: the hypothesis's DEFINING thesis is engaging 'ONLY on confirmed LARGE-magnitude directional regimes... rare, high-conviction legs... parabolic blow-offs and cascading crashes', FLAT otherwise, to remove noise dilution. On daily that was a rare ~20-day move (~4/yr). On 4H with entry_z default lowered to 1.0 (floor dropped to 0.75), it now engages on ~3.3-day, ~1-sigma moves 36-45x/yr with 37% time-in-market — an ORDINARY move, not a rare tail regime — so it no longer implements the stated core edge; it trades frequent moderate momentum. Per-trade capture collapsed from +4.15% (daily) to +0.85% (4H), consistent with a different, less-extreme phenomenon. Now-reliable sandbox shows a weak/dormant edge: +0.11% total, Sharpe -0.007, PF 1.01, avg_trade_return_pct -0.034%, trailing-365d -0.14%/trade. - CREDIT: the prior Layer-4 measurability blocker is genuinely RESOLVED. Moving to 4H raises the count from ~27 (unmeasurable) to 235 round trips over 6.6 years (~36/yr); the sandbox now places 72 trades from 36 entries with metrics_reliable=TRUE, top-3 concentration fell 60% -> 14.7%, kurtosis 4.7, holdout ~1.5 expected trades. Had the hypothesis specified 4H, this measurability profile would be acceptable — the block is the timeframe/edge mismatch, not the sample size. - The code itself is CORRECT and unchanged in mechanism from the iteration-1 version I verified (guarded z-score/lookback, n-1 sigma, expanding-vol gate, chandelier exit, risk-anchored leverage-capped sizing, _stdev instance-method fix). should_exit() still infers _side from the live z sign on restart (_side==0) — unreachable in backtest, only a live crash-restart risk. This fail is about hypothesis alignment, not implementation.

Verification Results

Analyst/optimizer/Research Lead: evaluate the holdout on pooled full-sample / walk-forward-OOS statistics rather than the 15-day window (as for the macro-TSMOM sibling). If the holdout gate is a hard un-waivable zero-trade auto-abandon, amend to a multi-asset pooled version or retire — don't burn repeated optimization runs on the same empty holdout. The optimizer will likely push entry_z UP (fewer trades), worsening the holdout.

Verification Results

Measurability is at/near the floor and the 15-day holdout will likely be EMPTY — the decisive call now belongs to the analyst/Research Lead, not QA. 74 round trips over 6.6 years (~11/yr) clears the ~30-trade unmeasurable line (2.7x the rejected iteration-1's 27) but sits below the ~100-trade floor at the default entry_z=1.25; the optimizer can reach 101 at the entry_z=1.0 bound, but its incentive (monotone frequency-vs-capture: 1.0->+1.09%, 2.0->+3.99%) pushes toward FEWER trades. The 15-day holdout expects ~0.46 trades, so it is structurally likely empty and may trip the zero-trade hard gate. Same profile as the macro-TSMOM sibling I passed (~16/yr, 103 total, waiver flagged), with a stronger reliable sandbox here (Sharpe 0.94, PF 4.23, +2.65%, positive 6/7 years, kurtosis 2.0, top-3 27.6%). Three iterations have mapped the frontier: a faithful single-name daily tail-regime tops out at ~74-101 trades.

Verification Results

Treat as 1-parameter (entry_z) for optimization/PBO; the chandelier multiple can be fixed.

Verification Results

The chandelier_atr tunable is effectively INERT (de facto 1-parameter strategy). With the 5-day window the z-sign-flip exit almost always fires before the 6x-ATR give-back, so chandelier 4.0 and 8.0 give identical results. The 'wide chandelier ride' is rarely the binding exit — exits are dominated by z-flip ('the tail regime is over'), itself a legitimate hypothesis exit, so this is not a code defect or missing mechanism (the chandelier is a correctly-implemented rarely-binding backstop). Lower effective DOF actually helps the deflated-Sharpe hurdle.

Verification Results

For live deployment, reconstruct _side/_extreme/_entry_atr from cache.positions_open() rather than the z sign.

Verification Results

should_exit() infers _side from the live z sign on restart (_side==0). Unreachable in backtest; only a live mid-position crash-restart risk.

Backtest Review

Genuinely novel, decorrelated convex-tail thesis (beta 0.009); good per-trade edge (PF 2.03, avg_trade_return_pct 2.23% well above fees); 2-param low DOF; tiny 2.25% drawdown; sharpe_ci_low 0.0516 (barely) positive

Backtest Review

Addresses the noise-dilution critique by standing flat most days

Backtest Review

Only 72 trades over 6.6y (~11/yr) — below the 100-trade measurability floor; with return_kurtosis 26.3 the Sharpe 0.675 is outlier-driven (a few days carry it)

Backtest Review

sharpe_ci_low 0.0516 sits right at zero — not distinguishable from no-skill

Backtest Review

Rare cadence breaks validation: the 15-day holdout expects ~0.5 trades (likely EMPTY hard-gate fail), WF-OOS windows sparse; 2-param sweep over ~72 fat-tail events would best-of-N fit and fail deflated Sharpe

Backtest Review

Tiny economic edge: +14.2% over 6.6y (~2%/yr), information_ratio -0.71 vs buy-hold BTC, 2025 negative

Outcome Summary

BtcTailRegimeConvexDirectionalLS took the opposite tack from the session's dead always-on trend followers: stay flat by default and engage only on confirmed large-magnitude directional regimes, riding crypto's fat-tailed blow-offs and crashes with a wide chandelier. The result had a genuinely clean shape — PF 2.03, 2.23% per trade, a 2.25% drawdown, beta 0.009 — but too little of it: +14.2% over 6.6 years on just 72 trades, a Sharpe of 0.675 whose CI touches zero, kurtosis 26.3, and underperforming buy-and-hold. The analyst abandoned it at backtest review because the rare cadence sits below the measurability floor and structurally breaks validation (likely-empty 15-day holdout, sparse OOS), so optimizing 2 parameters over a few blow-offs would best-of-N fit and fail deflated Sharpe. It never reached optimization, analysis, or risk review.

Outcome Summary

Removing noise-bar dilution by engaging only on rare extreme legs improves per-trade economics but trades away measurability — at ~11 trades/year the sample falls below the noise floor, the Sharpe is outlier-carried with a CI at zero, and the fixed 15-day holdout is likely empty, a cadence/validation mismatch optimization cannot fix.

Outcome Summary

The analyst abandoned it at backtest review: the idea is novel and decorrelated with a real per-trade edge, but it fires far too rarely to validate — only 72 trades (below the 100-trade floor), an outlier-driven Sharpe with a CI touching zero, and a cadence (~11/yr) that structurally breaks the framework, since the 15-day holdout expects ~0.5 trades (likely an empty hard-gate fail) and WF-OOS windows would be sparse. Optimizing 2 parameters over ~72 fat-tail events is best-of-N fitting that would fail deflated Sharpe.

Outcome Summary

A long-short, single-instrument convex directional strategy on BTCUSDT.BINANCE USD-M daily bars (2 parameters) that stayed flat by default and engaged only when a short 5-day trailing move exceeded a high z-score threshold with expanding volatility — riding confirmed large-magnitude directional regimes (parabolic bull legs, cascading crashes) with a wide chandelier stop, deliberately trading only the rare fat-tail legs to avoid noise-bar dilution.

Outcome Summary

The backtest (BTCUSDT.BINANCE 1D, 2409 data days) returned only +14.2% over 72 trades (~11/yr) with strong per-trade economics (profit factor 2.03, avg_trade_return_pct 2.23%), a tiny 2.25% drawdown, and genuine decorrelation (beta 0.009). But Sharpe was only 0.675 with a CI-low of 0.0516 (right at zero), return kurtosis was 26.3 (a few days carry it), and information ratio was -0.71 versus buy-and-hold BTC with 2025 negative.

Iteration History

Verification failed (Layer 4 — QA review): - STRUCTURALLY UNMEASURABLE — the tail-regime gate fires far too rarely to validate, so it cannot clear the significance/holdout gates and will burn optimization budget only to auto-abandon. Engaging only on confirmed large-|z| regimes with expanding vol yields ~4 trades per YEAR: 27 round trips over the full 6.6-year span and only 2 entries in the sandbox year (metrics_reliable=FALSE; the reported profit_factor 0.0 / win_rate 1.0 is a 2-winning-trade artifact, not a code defect). This is decisively below the ~100-trade measurability floor (L16) and the ~30-trade unmeasurable line (L26). Consequences the developer concedes: (a) the 15-day holdout expects ~0.17 trades, will be EMPTY, and trips the zero-trades hard gate -> auto-abandon; (b) OOS windows hold single-digit trades (meaningless Sharpes, uncontrollably wide CI on 27 obs); (c) top-3 trades are 60% of gross profit, so 27 observations cannot distinguish edge from luck. The developer states it 'can only be validated on a long window with pooled statistics, not by the standard 15-day-holdout gate.' Per L16/L26, reject at Layer 4 rather than spend the 3-phase optimization budget on a foreordained-empty holdout. - The code is CORRECT and faithfully implements the hypothesis — this fail is about sample size, not implementation. Verified: z=(log C[t]-log C[t-lookback])/(sigma_d*sqrt(lookback)) with past=_closes[0] (maxlen=lookback+1, guarded); sigma_d over vol_window, fast_sd over fast_vol_window, expanding=fast_sd>sigma_d; engagement requires |z|>=entry_z AND expanding, taking the side of z (WITH the move, not a fade); exits on z-sign flip or a wide chandelier from best close since entry; risk-anchored sizing capped with leverage. Polarity correct, symmetric long/short, all guards present, and the prior _stdev @staticmethod bug is fixed to a plain instance method. Full-sample stats (avg +4.15%/trade, PF 2.39, kurtosis 2.1, ~26% time-in-market, decorrelated) suggest the mechanism may be genuinely real — hence the block is measurability, not economics (unlike the negative-gross-edge ORB). - should_exit() infers _side from the live z sign on restart (_side==0) and re-seeds _extreme/_entry_atr from current values. Unreachable in backtest; only a live mid-position crash-restart risk. Moot given abandonment.
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.