Skip to content

View translation

BtcDayOfWeekSeasonalityLS

Hypotheses

BTC Day-of-Week Seasonality — Trade the Adaptively-Selected Strong/Weak WEEKDAY (Weekly Calendar Anomaly, Non-Trend / Non-Reversion), Low-Turnover Diversifier (BINANCE USD-M, Daily Bars, Long-Short, 3-Parameter)

Hypotheses

A LONG-SHORT, SINGLE-INSTRUMENT BTC probe that completes a systematic sweep of the CALENDAR-ANOMALY family — the one return source not yet in this factory's graveyard — across its three natural periods: monthly (pending turn-of-month sleeve), intraday (pending session-seasonality sleeve), and here WEEKLY (day-of-week). It is a deliberate, meta-learner-compliant departure from the trend family, which is now conclusively closed (L62: single-instrument trend survival 0.003; every non-momentum trend construction this session landed base Sharpe ≈0.4–0.7 and died, and the confluence 'edge-booster' thesis was falsified at 0.634; the TSMOM basket collapsed OOS to −3.3). The signal is purely the DAY OF WEEK — orthogonal to price path, trend, and mean-reversion — so it is a genuine diversifier to the portfolio's single promoted (momentum) strategy, not a correlated clone. Crypto plausibly exhibits weekday structure: reduced institutional participation and thinner liquidity on weekends vs weekday US/Europe hours, weekly derivatives expiry (Friday) positioning, and stablecoin/ETF flow cadences that concentrate on business days. It runs on the only artifact-free, non-fragile substrate this session found — single-instrument Binance USD-M pure-OHLCV (COIN-M times out, HL is history-capped, Deribit option data spans ~34 days, multi-instrument produces non-physical returns). To resist overfit the strong/weak weekdays are chosen ADAPTIVELY from a trailing window (not hardcoded), turnover is low (~1–2 one-day trades/week), and there is no optimizable rarity threshold. It avoids the graveyards: NOT trend (L62), NOT reversion/pairs/funding/factor (dead), NOT a liquidation/quarterly feed (L61). Exactly 3 tunable parameters: the trailing lookback for weekday selection, the number of weekdays traded per side, and the ATR stop multiple.

Hypotheses

The QA review asked for a full-history dry-run to rebut the claim that the +1.91% sandbox was a single-outlier artifact. I ran it on BTCUSDT.BINANCE 1-DAY, 2020-01-01..2026-08-06 (2410 bars), net of 0.05% taker per side. It does not rebut - it CONFIRMS, and I recommend ABANDON. 1) A weekday bias does exist in the raw data: Mon +42.9bps (t=+2.10), Wed +45.2bps (t=+2.44), Thu -20.5bps (t=-0.97), against an unconditional +14.3bps/day drift, and the sign is stable across sub-periods (Wed positive in 2020-21/2022-23/2024-26: +87.7/+9.0/+40.1bps; Mon +67.6/+8.2/+50.6). 2) But it is only harvestable with HINDSIGHT. A fixed long-Mon+Wed / short-Thu schedule chosen from the full sample returns +644% net (+0.262%/trade, Sharpe 0.99) - that is precisely the in-sample overfit this adaptive design was written to avoid, and it cannot be traded without knowing the answer in advance. 3) The ADAPTIVE estimator this strategy actually uses is negative in 40 of the 45 cells of its own declared parameter space (lookback_weeks 26/52/78/104/156 x days_per_side 1/2/3 x gate 0.00/0.35/0.75). The single best cell (lb=26, k=3, gate=0.75: +0.182%/trade, Sharpe 0.46) is the SHORTEST, noisiest lookback, whose picks are near-random (Wed chosen long only 47% of weeks, Thu short 50%) - a selection artifact, not the mechanism. I deliberately did NOT re-point the defaults at it. 4) The failure is not a selection failure. Long lookbacks DO recover the true days - at lb=156, k=2 the ranker picks Wed long 90% of weeks, Mon long 70%, Thu short 96% - and are STILL negative (-48.5% total, -0.006%/trade). Reason: forced top-k/bottom-k filling must short a second weekday that carries no bias, and that short's drag against BTC's +14bps/day drift exceeds the Mon/Wed edge. 5) The principled repair for (4) also fails. Dropping the forced top-k fill and trading a weekday only when its trailing demeaned mean clears a fee hurdle AND a t gate is negative in 27 of 27 configurations (lookback 52/104/156 x hurdle 10/20/30bps x t_min 0.0/1.0/1.5); best cell +0.007%/trade. A long-only variant peaks at +0.159%/trade with Sharpe 0.42 (vs +793% buy-and-hold) and flips sign on k (k=1 negative, k=2 positive) - another knife-edge. 6) Root cause is statistical power, not tuning: the demeaned weekday effect is ~30bps against a ~340bps daily sd, so a trailing window of 26-156 observations per weekday carries a 27-67bps standard error. No online estimator with a reachable lookback can resolve the effect; detecting it needs ~300 obs/weekday, i.e. essentially the whole sample. This is a structural impossibility for the adaptive design. Per the instruction not to regress passing layers, the code is UNCHANGED from the version that cleared Layers 1-3 (only a docstring block recording this measurement was added). No parameter change can fix a falsified mechanism, and shipping the one positive cell of a 45-cell grid would reproduce exactly the artifact QA flagged. This also closes the calendar-anomaly family on BTC: monthly (turn-of-month) was falsified by a full-parameter sweep, intraday session returned PF 0.10, and weekly is now falsified by a 72-configuration full-history measurement. Recommendation: ABANDON.

Hypotheses

Zero-edge single-name calendar seasonality, falsified by the developer's own measurement — not worth 2 hours of optimization. Sharpe is 0.115 with bootstrap CI [-0.69, 0.83] straddling zero, profit_factor 1.086, total_return 8.0% over 6.6 years (CAGR 0.91%), information_ratio -0.89, and avg_trade_return_pct 0.12% sits below the 0.15% Binance USD-M fee floor (L22), so the per-trade edge does not clear costs. The developer's docstring records a full-history 45-cell sweep in which the ADAPTIVE estimator is negative in 40 of 45 cells (the one positive cell is the shortest, noisiest lookback with near-random picks), and the principled fee-hurdle repair is negative in 27 of 27 configs; the stated root cause is statistical power, not tuning — the ~30bps weekday effect against a ~340bps daily sd gives a 27-67bps standard error on any 26-156-observation trailing window, so the estimator 'cannot converge faster than the effect decays, at ANY reachable lookback' — and it concludes the hypothesis is FALSIFIED on BTC with a recommendation to ABANDON. There is therefore no profitable parameter region for the 3-phase optimizer to find; re-pointing defaults at the single positive cell of 45 is exactly the selection artifact that would die at the DSR/PBO gate. Risk is contained (max_drawdown 8.3%, no liquidation), so this is not a blowup, but a ~zero-Sharpe, sub-fee-floor anomaly with no viable parameter region cannot be optimized into an edge. Failure pattern: no_edge/fee_edge single-name calendar seasonality (L22).

Implementation

Long/short BTC daily calendar-anomaly sleeve. Each completed daily bar's return is attributed to its weekday (derived from the bar timestamp, never a bar counter); a trailing per-weekday t-stat is maintained incrementally. Once per calendar week the top/bottom days_per_side weekdays by t become the long/short sets. The bar closing on weekday d decides the session of weekday d+1: enter at that close if the upcoming weekday is in the long (short) set and its t clears a multiple-testing bar, exit at the next close, hold through consecutive same-side weekdays, with an ATR stop and risk-capped sizing. ITERATION 3: a full-history dry-run now FALSIFIES the hypothesis on BTC - see rationale; recommend ABANDON.

Verification Results

CLEAN RESTART 2026-09-04 — this run's verdict history and learning records were removed and it was restarted from verification. Its previous abandonment came from the pipeline, not from the market: the Layer-2 harness mis-bound @staticmethod helpers (fixed), QA issued terminal performance verdicts on an unoptimized smoke test (removed — QA now judges correctness only), and sandbox timeouts came from backtest-slot starvation (fixed). The hypothesis and the strategy code are unchanged. Verify the code on its merits; performance is decided later by the full backtest and the optimizer.

Backtest Review

Genuinely orthogonal return source (weekly calendar), pure-OHLCV, no data-availability risk; contained drawdown 8.3%

Backtest Review

278 trades — adequate sample, so the flat result is a real no-edge read, and multiple-testing-aware t-gate is well-constructed

Backtest Review

Sharpe 0.115 with CI [-0.69, 0.83] straddling zero, PF 1.086, CAGR 0.91%, information_ratio -0.89 — no edge

Backtest Review

avg_trade_return_pct 0.12% is below the 0.15% Binance USD-M fee floor (L22) — per-trade edge does not clear costs

Backtest Review

Developer's own 45-cell sweep: adaptive estimator negative in 40/45 cells; fee-hurdle variant negative in 27/27; root cause is statistical power, not tuning — no reachable lookback converges

Backtest Review

Developer pre-registered ABANDON ('hypothesis is FALSIFIED on BTC')

Iteration History

Verification failed (Layer 2 — synthetic scenarios): Parameters used: ['_risk_frac', '_atr_period', '_gate_scale', '_param_bounds', 'atr_stop_mult', 'days_per_side', '_max_hold_days', '_min_stop_frac', 'lookback_weeks', '_max_gross_frac', '_min_sample_frac'] Check that __init__ sets all attributes from self.parameters.get(). - steady_uptrend: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000) - steady_downtrend: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000) - flat_ranging: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000) - volatility_spike: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000) - zero_volume: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000) - price_gap: TypeError: FactoryStrategy._bar_ts() takes 1 positional argument but 2 were given (bar timestamp: 1735690740000)

Iteration History

Verification failed (Layer 4 — QA review): - No evidence of a real, robust edge -- the marginally-positive sandbox is a single-outlier artifact, not the consistent weekday bias the mechanism claims (checklist #7, L6/L7). Sandbox: total_return +1.91%, avg_trade_return_pct +0.332% (nominally above the floor) BUT Sharpe 0.20 with CI [-1.58, 1.66] straddling zero, PF 1.15, and skew 2.92 / kurtosis 43 over 60 trades. The largest win (2769) is ~6x the average win (453); backing out that one trade leaves the book roughly breakeven-to-negative (avg_win 453 < avg_loss 484, win_rate 0.55). So the +1.91% rests on essentially ONE outlier -- for a day-of-week seasonality the edge should be a consistent small per-weekday bias, and a one-trade-driven, kurtosis-43 result is the signature of noise, not the mechanism working. No full-history dry-run was provided to rebut this. Decisively, this is the THIRD calendar-anomaly probe on BTC by the same developer, and the family is dead here: the monthly (turn-of-month) probe was falsified by a full-parameter-space measurement (no boundary shape, recent regime negative across all windows) and the intraday session probe returned PF 0.10 (anti-predictive). This weekly probe does not establish that day-of-week is any different.

Iteration History

Verification failed (Layer 4 — QA review): - The hypothesis's core edge is empirically FALSIFIED on BTC, and the developer's own iteration-3 full-history dry-run (BTCUSDT.BINANCE 1-DAY, 2020-01-01..2026-08-06, 2410 bars, net of 0.05% taker/side) settles it — the developer correctly recommends ABANDON. (1) The ADAPTIVE estimator this strategy actually uses is negative in 40 of the 45 cells of its own declared parameter space (lookback_weeks 26/52/78/104/156 x days_per_side 1/2/3 x gate 0.00/0.35/0.75); the single positive cell (lb=26,k=3,gate=0.75) is the SHORTEST, noisiest lookback whose weekday picks are near-random (Wed selected long only 47% of weeks, Thu short 50%) — a selection artifact, not the mechanism. (2) Longer lookbacks that DO recover the true weekday bias (lb=156,k=2 picks Wed long 90% / Thu short 96%) are still negative (-48.5% total), because forced top-k/bottom-k filling must short a second, unbiased weekday against BTC's +14.3bps/day drift, and that drag exceeds the Mon/Wed edge. (3) The principled repair (drop the forced fill, trade only weekdays whose trailing demeaned mean clears a fee hurdle plus a t gate) is negative in 27 of 27 configurations. (4) Root cause is statistical power, not tuning: the demeaned weekday effect is ~30bps against a ~340bps daily sd, so 26-156 observations/weekday carry a 27-67bps standard error — resolving the effect needs ~300 obs/weekday (essentially the whole sample), a structural impossibility for any reachable adaptive lookback. The bias is real but only harvestable with HINDSIGHT (a full-sample-fixed long-Mon+Wed/short-Thu schedule returns +644%, Sharpe 0.99) — precisely the overfit this adaptive design was written to avoid. - The positive Layer-3 sandbox (+1.91%, avg_trade_return_pct 0.332%) is an outlier artifact, not evidence of edge, and does not survive the full-history measurement above. It is a single 362-day slice with Sharpe only 0.20, sharpe_ci_low -1.58 (CI spans zero and deeply negative), return_skew 2.92 and return_kurtosis 43.0 — one largest_win of 2769 on 60 trades dominates the record — with impact_cost_pct 17.3% and capacity only $3.3M. My bar requires the positive sandbox to be corroborated by (not contradicted by) the full history; here the developer's full history is net-negative across essentially the entire parameter space.
Strategy report

Backtest and paper results are hypothetical. Trading involves risk of loss.