Skip to content

View translation

BtcVolatilityTargetedTrendLS

Hypotheses

BTC Volatility-Targeted Trend Exposure, Long-Short (Single-Instrument BTCUSDT.BINANCE Perp, Daily Bars — Directional Sign from a Medium Trend, Position Size Scaled INVERSE to Realized Volatility, Drawdown-Controlled, 2-Parameter)

Hypotheses

A LONG-SHORT, single-instrument, pure-OHLCV volatility-TARGETED trend strategy on BTCUSDT.BINANCE USD-M perpetual — a materially different CONSTRUCTION and edge source from everything in my pending book (all dual/triple/macro momentum-CONFLUENCE and relative-value, which key on ENTRY selection). Here the edge is not clever entry timing but VOLATILITY TIMING: the position DIRECTION is a simple medium-horizon trend sign, and the position SIZE is scaled INVERSE to realized volatility toward a constant risk target — the documented managed-futures technique that improves risk-adjusted returns and, critically, CONTROLS DRAWDOWN by automatically de-risking in turbulent regimes and up-sizing in calm trends. It directly targets the failure dimension now killing strategies: the relative-momentum ideas died not only on fee-marginal edge but on CATASTROPHIC drawdown (XRP/ETH: 65% DD, promotion-disqualifying); vol-targeting is the standard, principled answer to exactly that. It reuses the one asset with demonstrated directional tradability (BTC — every alt died 'no edge'), stays pure OHLCV (the only reliably measurable, coverage-complete data), single-instrument, developer-safe with standard incremental indicators, and DELIBERATELY 2-PARAMETER (trend lookback + target volatility) to resist overfit. A ~50-day trend flips roughly monthly → ~80-150 directional turns over multi-year history (measurable), while the vol-scaling continuously adjusts size. It fills the under-target LONG-SHORT bucket (13.6% vs 86.4% long-only) and adds a genuinely distinct strategy class (risk-managed trend / volatility timing) to a momentum-saturated portfolio.

Hypotheses

My recommendation is ABANDON, and this artifact deliberately changes no trading behaviour. In iteration 2 I pre-committed that if the enlarged sample still showed no edge then the claim under test — that size management alone lifts risk-adjusted returns on a SINGLE asset — is falsified and a one-market version should be abandoned rather than re-clocked a third time. Layer 4 returned that verdict and a stronger one: 14 trades over 362 days, directional turns projecting to ~50-60 over six years (below the ~100 floor), and the entire +20.24% carried by one $24,250 win that exceeds total net profit, with a Sharpe CI straddling zero. QA's explanation of why my trigger-rate prediction failed is also correct and structural rather than a bug: I assumed flip frequency scales as 1/L, but BTC trends autocorrelate, so shortening the clock buys far fewer independent turns than the arithmetic implies. The obvious 'fix' — re-clock 20 to 10 days — is precisely what I promised not to do: it degrades the direction reading toward noise, multiplies fee drag on a construction already paying ~3%/year, and pads the sample with resize-style observations that are not independent. That is best-of-N fitting dressed as an iteration, and it would also invalidate the two-parameter design that is this hypothesis's main defence. The tail dependence is likewise not a defect to repair: few large winners carrying the result IS the trend-following payoff, and truncating it would remove the edge rather than measure it — the honest reading is that 14 trades cannot separate that profile from luck. So the only change is diagnostic, and it exists to make the verdict auditable rather than reconstructed by hand: on_stop now prints DIRECTIONAL TURNS PER YEAR (turns normalized by bar count — the statistic the measurability floor is actually about, as opposed to raw trade count which mixes in resizes) and a concentration line giving closed-position count, net P&L, the largest single win, and net P&L EXCLUDING that win. If that last figure is negative, the run is one-trade-dependent by construction, which is exactly the test QA performed manually. Neither addition touches an order, a size, or a threshold, and both run once at stop. For the record on where this technique should go: classic vol-targeting earns its statistical power from breadth — diversifying a slow signal across dozens of markets so independent observations come from instrument count rather than from clocking one asset faster. That is a different hypothesis worth proposing on its own terms, not an iteration of this one.

Hypotheses

Promotion-disqualifying drawdown with no significant edge — not worth 2 hours of optimization. max_drawdown is 52.7% (CI to 82.3%), past the L19 50% hard-abandon line, on a strategy whose stated purpose was drawdown control — single-asset vol-targeting did not deliver it. There is no significant edge: Sharpe 0.446 with bootstrap CI [-0.31, 1.14] straddling zero, profit_factor 1.07, information_ratio -0.54, and the +104.5% headline is HALF open-position unrealized (end_unrealized_pct 50.9), so the realized track record is only ~53% and this is not a validated realized edge. The result is bull-cycle-dependent and recently catastrophic: 2020 +94% and 2023 +61% carry it while 2024 is -25.0% and 2025 -28.6% (rolling Sharpe to -5.4), so the last-20% holdout sits in a deep loss. The developer pre-registered ABANDON on exactly this evidence: 'classic vol-targeting earns its power by diversifying a slow signal across dozens of markets, and a one-market version should be ABANDONED rather than re-clocked a third time' — a single-asset/sample-size verdict, not a bug. The proper vol-targeting construction is multi-market, which is a different hypothesis (not tunable here), and no parameter change fixes a 52.7% drawdown on one asset. Unlike the dual-TF momentum variants advanced to optimization this batch (Sharpe ~1.0-1.1, DD <22%, positive holdout), this fails the drawdown hard gate and the significance test. Failure pattern: risk_reject/no_edge single-asset vol-targeted trend, >50% DD (L19).

Implementation

Long/short volatility-targeted trend exposure on the BTCUSDT.BINANCE USD-M perpetual, 1-DAY bars, pure OHLCV. Direction is the sign of a sigma-normalized medium-horizon trend reading — (close - close[-L]) / (ATR x sqrt(L)) — with a locked 0.25-sigma dead zone; size is the edge: target notional = equity x (target_vol / realized 30-day annualized vol), capped at leverage x 0.9 of equity, with self.config.leverage read directly in the sizing path so the declared margin headroom is genuinely consumed by calm-regime up-sizing. Exits are structural and untunable: a trend flip to the opposite band (the flip is the stop in this construction — risk control is size, not exit timing) and a volatility resize when held notional drifts more than resize_band from the current target, which converts entry-time targeting into continuous targeting using only the base hooks. Two tunables: trend_lookback and target_vol. Iteration 3 changes NO trading behaviour — the only additions are on_stop diagnostics reporting directional turns per year and P&L concentration (net P&L excluding the largest single win).

Verification Results

CLEAN RESTART 2026-09-04 — this run's verdict history and learning records were removed and it was restarted from verification. Its previous abandonment came from the pipeline, not from the market: the Layer-2 harness mis-bound @staticmethod helpers (fixed), QA issued terminal performance verdicts on an unoptimized smoke test (removed — QA now judges correctness only), and sandbox timeouts came from backtest-slot starvation (fixed). The hypothesis and the strategy code are unchanged. Verify the code on its merits; performance is decided later by the full backtest and the optimizer.

Verification Results

Strategy doesn't achieve its core purpose: sold as drawdown-controlled vol-targeting, it posts 51.9% max drawdown (CI to 83.5%) with a marginal, zero-spanning Sharpe (0.40) over 112 trades. Leverage-2.0 up-sizing in calm regimes (cap 1.8x) magnifies losses on reversals. Faithful implementation, so a thesis-validity finding for the analyst — single-asset vol-targeting doesn't reproduce the drawdown control the technique gets from cross-market diversification. Abandon the single-asset form; don't re-clock the lookback; flag the DD to the Risk Officer.

Backtest Review

Clean construction, 114 trades (measurable), standard incremental indicators, no data risk

Backtest Review

Positive alpha (+0.15) and works in trending bull cycles (2020 +94%, 2023 +61%)

Backtest Review

max_drawdown 52.7% (CI to 82.3%) — past the L19 50% hard-abandon line, on a strategy whose stated purpose was drawdown control

Backtest Review

Sharpe 0.446 with CI [-0.31, 1.14] straddling zero, profit_factor 1.07, information_ratio -0.54 — no significant edge

Backtest Review

Headline +104.5% is half open-position unrealized (end_unrealized_pct 50.9); realized track record ~53% is materially weaker

Backtest Review

Catastrophic recent years (2024 -25%, 2025 -28.6%), rolling Sharpe to -5.4 — the holdout window is deeply negative

Backtest Review

Developer pre-registered ABANDON: single-asset vol-targeting is falsified (the technique needs multi-market breadth)

Iteration History

Verification failed (Layer 4 — QA review): - UNMEASURABLE SAMPLE, STRUCTURAL TO THE CONSTRUCTION -- a single-asset, daily, medium-horizon trend sign produces too few independent bets to validate. Sandbox produced 7 trades over 362 days. This is not a bug: entries have no fresh-cross (stay-in-market by design), so trades = trend flips + vol resizes, and a 50-day trend with a 0.25-sigma dead zone flips rarely (avg_holding 37 days) while the 60% resize band is deliberately wide -- so ~7/year is the inherent rate. Projected over a multi-year span that is ~45-60 directional turns, below the ~100 measurability floor (L16), and short of the hypothesis's own '~80-150 over multi-year' claim. At n=7 the result is pure noise: total_return +3.50% is carried entirely by 2 winners of 7 (avg_win $11,956 vs avg_loss $4,093, largest_win $13,396), win_rate 0.286, long_win_rate 0/3, Sharpe 0.209 with CI [-1.58, 1.68] straddling zero, return_kurtosis 12.1. avg_trade_return_pct 1.43% cannot be read as per-trade edge on a 2-win sample. This is the same measurability wall the confluence re-timings hit, here for a structural reason -- classic managed-futures vol-targeting earns statistical power by diversifying a slow signal across DOZENS of markets; one BTC daily trend sign cannot supply a measurable directional sample.

Iteration History

Verification failed (Layer 4 — QA review): - STILL SUB-MEASURABLE AND ONE-TRADE-DEPENDENT -- the re-clock lifted the count but not enough, and the positive result is a single win. total_trades rose only 7 -> 14 over 362 days, well short of the developer's own predicted ~35 round trips: the trend_lookback 50 -> 20 change assumed flip frequency scales as 1/L, but BTC trends autocorrelate, so 20-day trends still flip infrequently. Over a ~6-year span 14/year is ~84 total trades, and the DIRECTIONAL TURNS -- which the developer correctly identifies as the ONLY independent observations (resizes merely re-establish the same bet) -- are fewer still, projecting to ~50-60, below the ~100 measurability floor (L16). Worse, the +20.24% total_return is carried by ONE trade: largest_win $24,250 exceeds the entire net profit (avg_trade $1,442 x 14 ~= $20,187), so removing the single biggest win makes the sandbox year net-negative. win_rate 0.429, long_win_rate 1/6, Sharpe 0.643 with CI [-1.05, 2.13] straddling zero, return_kurtosis 9.79. This is the identical tail-luck signature as iteration 1 (2 winners of 7); the sample simply has too few independent winning observations to distinguish edge from luck.

Iteration History

Verification failed (Layer 4 — QA review): - CONFIRMED SUB-MEASURABLE AND ONE-TRADE-DEPENDENT -- unchanged from iter2 by design, and the developer's own recommendation is ABANDON. This iteration changes NO trading behaviour (identical parameters, so the sandbox is byte-identical to iter2: 14 trades, total_return +20.24%, Sharpe 0.643 with CI [-1.05, 2.13] straddling zero, PF 1.544). The result remains what iter2 established: directional turns project to ~50-60 over a multi-year span (below the ~100 measurability floor, L16), and the entire +20.24% is carried by a single largest_win $24,250 that exceeds total net profit (avg_trade $1,442 x 14 ~= $20,187) -- remove it and the year is net-negative. At 14 trades this tail-driven profile cannot be distinguished from luck, which is a sample-size verdict, and it is structural to a single-asset daily slow-trend construction (BTC trends autocorrelate, so re-clocking faster buys non-independent turns and fee drag, not statistical power).
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.