Skip to content

View original

BtcFourHourVolumeConfirmedDonchianBreakoutLongDailyRegime

Hypotheses

BTC 4H Volume-Confirmed Breakout Long with Daily Regime Filter

Hypotheses

A long-only single-instrument breakout strategy on BTCUSDT perpetual futures using 4-hour bars and OHLCV-only data, with a daily-bar regime filter. Architecturally MIRRORS the proven BNB 4H Volume-Confirmed Breakout (Sharpe 2.03, paper_stage), applied to BTC as the third sensitivity-validation test point of this construction class (BNB ✓ proven, SOL ⏳ in pipeline, LINK ⏳ in pipeline, BTC proposed here). BTC is the SAFEST asset for this architecture-transfer because: (a) BTC has empirically validated trend persistence at multiple timeframes (BTC Daily Multi-Week Trend Continuation = pending, BTC Spot Drawdown Accumulation = Sharpe 3.33 paper_stage), (b) BTC has the deepest liquidity and cleanest 4H microstructure of any crypto perp — institutional flows produce the smoothest momentum continuation signals, (c) BTC's correlation to BNB is ~0.75-0.85 — high enough that the architecture should transfer, with the additional advantage that BTC has CLEANER trend dynamics than BNB (BNB Golden Cross failed because BNB has stronger mean-reversion at daily; BTC has the OPPOSITE empirical character at daily, so 4H momentum should be even cleaner than BNB's). Importantly addresses the recently-documented 'auto-replication failure pattern' by: (i) requiring sensitivity_passed=true with cliff_count=0 as a HARD GATE before optimization, (ii) selecting BTC based on existing empirical evidence of clean trend behavior across multiple timeframes (NOT a thesis-only argument like BNB Golden Cross). Existing BTC strategies in pipeline (BtcDailyMultiWeekTrendContinuationLong + BtcSpotDrawdownAccumulationLong) use daily bars + sustained momentum / drawdown DCA — this proposal uses 4H bars + intraday volume-confirmed breakouts: orthogonal in BOTH timeframe AND mechanism, providing genuine portfolio diversification with minimal signal correlation to existing BTC exposure. Single-dominant-filter design (200-day SMA daily regime + 20-bar 4H Donchian high + volume confirmation), explicitly NOT a multi-condition AND-gate. Calibrated for ~25-45 entries/year — well above the sparsity-failure threshold.

Hypotheses

Iteration 2 addresses optimization over-fit selection, not a code defect: the feedback confirms the mechanism is genuinely sound (walk-forward is_overfitted=FALSE, all 3 OOS windows positive, 0 cliffs, positive holdout) but best-of-225 grid selection chose a Sharpe-maximizing config (PBO 0.73, holdout ratio 0.25) by thinning the book to 86 fat-tailed trades. Since the strategy already reached optimization, every earlier verification layer (static, synthetic, sandbox) passed — so the signal/exit/sizing code is kept byte-identical to avoid regressing a green layer. The single lever the developer controls to steer selection toward the robust base region is the config defaults, which the sensitivity/walk-forward optimizer varies around. All defaults are pinned dead-center in the feedback's requested robust ranges (exit_period 10 in 9-12, donchian_period 20 in 18-24, vol_mult 1.5 in 1.4-1.7, stop_pct 0.08 in 0.07-0.09, regime_sma 200 in 180-220). This is the ~101-trade base config that showed Sharpe 1.29, sharpe_ci_low +0.28, positive in 6/7 years and maxDD 7.6% — a config near these defaults targets holdout ratio >=0.70 and deflated_sharpe >=0.95 with PBO <=0.5, trading a lower headline Sharpe (~1.3-1.6) for forward-robust, defensible generalization rather than the over-fit Sharpe-1.90 pick.

Hypotheses

Failed deflated Sharpe on the final attempt (2 of 2): DSR=0.4926 (vs 0.95 bar), is_significant=FALSE, with the optimized Sharpe 1.5532 AT/BELOW the expected-max best-of-225 luck bar of 1.5629 and PBO=0.7249 (>0.5) — after multiple-testing correction the selected config is statistically indistinguishable from best-of-N noise. This is the textbook PSR-vs-DSR trap: probabilistic_sharpe 0.9985 ignores the 225-trial count while DSR, which corrects for it, rejects. The passing forward gates (walk-forward is_overfitted=false with avg IS 2.71 -> OOS 1.47; holdout ratio 1.041; sensitivity 0 cliffs; sharpe_ci_low 0.44>0) measure CONSISTENCY, not significance, and cannot override the deflation failure — a consistent, non-overfit ~1.3-1.5 Sharpe that sits below the luck bar is still a sub-significant edge. Not iterate: attempt 2 of 2 is exhausted and the sensitivity surface is a flat plateau (whole grid ~1.0-1.44 Sharpe, 0 cliffs) — the optimizer is already on its plateau with no robust region above the luck bar to tune toward, so a third sweep re-selects the same sub-significant config. Not revise_hypothesis: BTC is not a dead target and the volume-confirmed Donchian breakout mechanism already has its promoted/paper instance on BNB (Sharpe 2.03) — this is the same mechanism ported to BTC proving under-powered (asset/variant-selectivity failing deflation), not a proven mechanism stranded on a structurally-dead instrument. FAILURE PATTERN: the FourHourVolumeConfirmedBreakout architecture does NOT auto-transfer from its promoted BNB instance to every major — on BTC the genuine but modest trend-breakout edge (clean 0-cliff sensitivity, non-overfit walk-forward, passing holdout) cannot clear best-of-225 deflation: the selected Sharpe 1.55 lands right at the 1.56 luck bar, DSR 0.49, PBO 0.72, IR -0.61 vs its basket benchmark. This is the identical signature to the abandoned SUI arm of this family — a validated family + clean sensitivity grid + passing holdout do not rescue an asset whose underlying edge is too modest to be distinguished from selection noise; BNB remains the instrument where this mechanism clears the bar.

Implementation

BTCUSDT 4H long-only volume-confirmed Donchian breakout with a 200-day daily-SMA regime filter. Enters long when the daily close is above its 200-day SMA (bull regime), the 4H close breaks above the prior 20-bar Donchian high, and current 4H volume exceeds 1.5x its 20-bar average. Exits on a trailing 10-bar Donchian-low break or an 8% hard stop. Signal is a continuous, graded breakout distance in ATR units. Capital-relative risk-based sizing (2% risk per trade, 50% notional cap), leverage 1.0.

Backtest Review

Clean trend-following signature that faithfully implements the hypothesis: 37.9% win rate with strongly asymmetric payoff (avg_win $3,621 vs avg_loss $1,051), PF 2.10, positive skew +7.37, Sortino 5.36

Backtest Review

Low risk profile — max_drawdown 7.84%, Calmar 10.8, exposure only 28.5% (idle capital = scaling capacity)

Backtest Review

Daily 200-SMA regime filter works as designed: flat through the 2022 bear (no entries), positive alpha +0.04, PSR 0.998

Backtest Review

Sibling-validated architecture (BNB 4H Volume-Confirmed Breakout, Sharpe 2.03); 95 entries submitted = 95 signaled, no size/notional drops

Backtest Review

Modest fee drag (commission 2.9% of gross, impact 1.2%)

Backtest Review

Trade count light at ~95 over 6.5 years (~15/yr, below the 25-45/yr target) — the post-optimization deflated-Sharpe / multiple-testing gate will be the real hurdle

Backtest Review

Recent-regime decay: 2025 annual -0.87%, rolling Sharpe went negative mid-2025 — holdout must confirm the edge still works forward

Backtest Review

High kurtosis (72) and skew (+7.37): returns are carried by a few outlier days (Jan/Feb 2021) — normal for breakout but concentration risk

Backtest Review

information_ratio -0.61 — lags its benchmark, though the comparison is loose for a low-exposure timing strategy

Analysis

Forward-consistency gates all pass: walk-forward is_overfitted=false (avg IS 2.71 -> avg OOS 1.47, ratio 1.84 < 3.0), holdout passed (ratio 1.041, holdout_sharpe 1.53 vs WF-OOS 1.47).

Analysis

Clean sensitivity surface: 0 cliffs across all 11 parameters; the vol_mult x regime_sma heatmap is a smooth plateau, and sharpe_ci_low 0.4443 is > 0.

Analysis

Correct, low-drawdown construction: max_drawdown ~6-8%, profit_factor ~2.0, sortino ~4.7, capacity is large ($427M), fees are modest (commission 3.43% of gross).

Analysis

Fails deflated Sharpe decisively: DSR 0.4926 << 0.95, is_significant=false, and the optimized Sharpe 1.5532 sits AT/BELOW the expected-max best-of-225 luck bar of 1.5629 — indistinguishable from selection noise.

Analysis

PBO 0.7249 > 0.5 — the parameter selection is more likely than not overfit.

Analysis

PSR 0.9985 vs DSR 0.4926 is the classic multiple-testing trap: the high PSR ignores the 225-trial count that DSR corrects for and rejects.

Analysis

Negative information_ratio (-0.61) vs its (meaningful) basket benchmark; edge is outlier-carried (kurtosis 31.5, skew 4.7) and decaying recently (optimized 2025 +1.67%, 2026 -1.53%).

Analysis

Sensitivity is a flat plateau (~1.0-1.44 Sharpe everywhere) — no robust region above the luck bar exists to tune toward.

Analysis

Do NOT promote yet, but this is a genuine, non-overfit edge worth one refinement pass — the mechanism is sound (walk-forward is_overfitted=FALSE, all 3 OOS windows positive, 0 cliffs, Sharpe above the luck bar, positive holdout). The blocker is that best-of-225 selection over-fit a Sharpe-maximizing config (PBO 0.73, DSR 0.73, holdout ratio 0.25). The optimizer pushed exit_period 10->6, donchian_period 20->30, vol_mult 1.5->1.89, stop_pct 0.08->0.054, thinning the book to 86 trades and softening in the recent holdout window. SPECIFIC CHANGES: (1) Re-run optimization with a NARROWER, conservative parameter search centered on the robust base region rather than the full grid — e.g. exit_period 9-12, donchian_period 18-24, vol_mult 1.4-1.7, stop_pct 0.07-0.09, regime_sma 180-220. A tighter search reduces the multiple-testing penalty (effective n_trials) that is suppressing DSR and lowers PBO. (2) Select for HOLDOUT GENERALIZATION and DSR, not in-sample Sharpe: the default config already showed Sharpe 1.29 with sharpe_ci_low +0.28, positive in 6/7 years, and max_DD 7.6% — a config near those defaults likely passes the holdout ratio (>=0.70) where the Sharpe-1.90 config failed (0.25). (3) Keep trade count higher (the base config's ~101 trades vs the optimized 86) to reduce per-trade fat-tail dependence (optimized kurtosis 30.9) and lift DSR. TARGET for the next pass: holdout ratio >= 0.70 AND deflated_sharpe >= 0.95 with PBO <= 0.5, accepting a lower headline Sharpe (~1.3-1.6) in exchange for forward robustness. The edge is real and this is the best BTC-arm of the volume-breakout family — the goal is a defensible, generalizable config, not the Sharpe-maximizing one.

Outcome Summary

This strategy took the factory's proven BNB volume-confirmed Donchian breakout architecture and ported it to BTC, arguing BTC's deep liquidity and clean trend dynamics made it the safest transfer target, and it deliberately built in a hard sensitivity gate to avoid the auto-replication failure pattern. It delivered the cleanest profile of its cohort: a strong low-drawdown base (Sharpe 1.33, profit factor 2.10, 7.8% max drawdown) and an optimized config that passed every consistency check — a non-overfit walk-forward, a passing holdout, zero sensitivity cliffs, and a positive Sharpe CI-low. But it still failed deflation: its optimized Sharpe of 1.55 landed right at the 1.56 best-of-225 luck bar, with DSR 0.49 and PBO 0.72, and the flat sensitivity plateau offered no region above the bar to reach for. The analyst abandoned it on attempt 2 of 2, concluding the breakout architecture does not auto-transfer from its promoted BNB instance to every major — on BTC the edge was real and stable but too modest to be distinguished from selection noise, so BNB remains the only instrument where this mechanism clears the bar.

Outcome Summary

A validated architecture with a promoted sibling does not auto-transfer to every major asset — passing walk-forward, holdout, and sensitivity gates prove the edge is consistent and non-overfit but not that it is statistically significant, so a genuine-but-modest edge whose optimized Sharpe merely matches the best-of-N luck bar (DSR ~0.49, PBO ~0.72) must still be rejected.

Outcome Summary

The analyst abandoned it on the final attempt (2 of 2) because, despite clean consistency gates, it failed the deflated Sharpe test decisively (DSR 0.493 vs 0.95, not significant) with the optimized Sharpe 1.553 sitting at/below the best-of-225 luck bar of 1.563 and PBO 0.725 — after multiple-testing correction the config was statistically indistinguishable from selection noise, and the sensitivity surface was a flat ~1.0-1.44 plateau with no robust region above the luck bar to tune toward.

Outcome Summary

A long-only 4H volume-confirmed Donchian breakout on BTCUSDT perp — entering when price broke the 20-bar Donchian high with above-average volume while a daily 200-SMA regime filter confirmed a bull market — ported directly from the factory's proven BNB 4H Volume-Confirmed Breakout architecture (Sharpe 2.03, paper stage) as an architecture-transfer test on BTC.

Outcome Summary

The base backtest was strong and low-risk: Sharpe 1.33, +79.5% return, profit factor 2.10, Sortino 5.36, Calmar 10.8, and only 7.84% max drawdown over 95 trades (37.9% win rate with asymmetric payoffs). Optimization produced Sharpe 1.553, +82.7% return, 6.2% drawdown, and profit factor ~2.0, and it uniquely passed the consistency gates — non-overfit walk-forward (avg IS 2.71 → avg OOS 1.47), passing holdout (ratio 1.041, Sharpe 1.53), zero sensitivity cliffs, and a positive Sharpe CI-low (0.444).
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.