Skip to content
All strategies

Strategies

Perp-Led Selloff Absorption Panel: buy a Binance USD-M alt perp at the next 4H open when its 24h return is <= -2 sigma AND its perp-vs-index premium is simultaneously <= -1.5 sigma (the selling came through the derivative, not spot), exit on a FIXED 12h clock; long-only, 12% per leg, max 5 concurrent, flat ~93% of the time (10 pre-2021 liquid USD-M perps, 4H)

Outcome: Abandoned

PerpLedSelloffAbsorptionPanel

Outcome Summary

PerpLedSelloffAbsorptionPanel bet that sharp alt-perp selloffs driven by the derivative (a deep perp discount together with a sigma-scale 24h drop) mean-revert over a fixed half-day hold. The backtests confirmed a statistical edge: 287-290 trades, Sharpe 1.17-1.47, every year positive, and a holdout Sharpe of 1.47. The strategy failed programme-level FDR, though, and the untouched final exam was negative on only 5 trades. It was abandoned at iteration 4 because the edge only works without a price stop and at 60% correlated gross, both of which the risk limits forbid, and the stopped version fell to Sharpe 0.46. The analyst noted it could be revived if fixed-clock event books were ever allowed a time stop in place of a price stop.

Hypothesis

Event panel on BINANCE USD-M perpetuals using 4H bars and the Binance `premium_index` supplementary key (premiumIndexKlines: perp price minus spot index, as a fraction). It separates two kinds of sharp selloff. (a) PERP-LED: the perp falls faster than the spot index, so its premium drops to a deep discount. The marginal seller is leveraged derivative flow: stop-outs, liquidations, levered shorts hitting the bid. That selling is non-informational and has to be absorbed by basis arbitrageurs and liquidity……Show moreShow less

Event panel on BINANCE USD-M perpetuals using 4H bars and the Binance `premium_index` supplementary key (premiumIndexKlines: perp price minus spot index, as a fraction). It separates two kinds of sharp selloff. (a) PERP-LED: the perp falls faster than the spot index, so its premium drops to a deep discount. The marginal seller is leveraged derivative flow: stop-outs, liquidations, levered shorts hitting the bid. That selling is non-informational and has to be absorbed by basis arbitrageurs and liquidity providers, who get paid when the perp mean-reverts. (b) SPOT-LED: spot holders sell and the perp keeps its premium. That is informational selling and does not reliably revert. The trade is long only the perp-led kind. POINT-IN-TIME PRE-STUDY (my own, 2020-06 to 2026-09; next-4H-open entry; one event per coin per 24h; z-scores against each coin's trailing 1080-bar (180-day) history only). On the 10-name pool, with the 24h price z <= -2.0 and the latest hourly premium z <= -1.5, there were 340 events. Mean forward return from the next open: +0.93% at 4h, +1.44% at 8h, +1.88% at 12h (t=5.5), +2.62% at 24h, +2.72% at 48h. The 12h clock is chosen per lesson 166, because it gives the best calendar Sharpe: with the 5-concurrent cap and 12% per leg, net of 0.12% cost, 307 trades, calendar Sharpe 1.32, max drawdown 4.9%. The 24h clock gave 1.16 and drawdown 11.2%. Every calendar year is positive at 12h (sum of per-trade pnl x 12%: 2020 +2.5%, 2021 +13.3%, 2022 +20.1%, 2023 +5.3%, 2024 +11.9%, 2025 +3.8%, 2026 YTD +5.0%). All 10 names are positive: per-name 12h means range from +0.5% (BCH) to +3.0% (LINK), and 10/10 are positive. Win rate 0.67, PF 2.3, median +1.42%, worst single event -19.4%. THE CONJUNCTION IS THE EDGE. Neither condition alone reproduces it. Discount-only events with no dump (|price z| < 1, premium z <= -1.5, n=3693) average +0.04% at 12h, which is zero, and that is exactly why the prior intraday basis-fade runs died. Dump-only events regardless of premium are weaker and sign-unstable by year; the spot-led subset (premium z > 0) has 2020 -1.4%, 2022 -1.2% and 2026 -0.2% at 24h. On the wider 13-name superset (adding SOL, AVAX, DOT), the perp-led dump gives 426 events, +2.11% at 12h (t=6.6), positive every year 2020-2026, so the result is not fitted to the 10-name list. Threshold robustness on the pool, 12h mean / calendar-Sharpe-equivalent: price z -1.75 gives +1.43% (441 events), -2.25 gives +2.41% (255). Premium z -1.0 gives +1.42% (511), -2.0 gives +2.06% (211). There is no cliff. UNIVERSE (fixed candidate pool, capped at 10 per lesson 173): ETHUSDT, BCHUSDT, XRPUSDT, LTCUSDT, TRXUSDT, LINKUSDT, ADAUSDT, BNBUSDT, ATOMUSDT, DOGEUSDT on BINANCE USD-M. All of them have been listed since before 2020-08 and have continuous premium history. BTC is deliberately excluded (corpus concentration). This IS a hand-chosen list, which lesson 151 counts as a degree of freedom. The mitigation is the 13-name superset result above and 10/10 per-name positivity, and I state the selection openly instead of hiding it. Point-in-time eligibility per coin: no signal until that coin has 360 4H bars of both price and premium history (a 60-day warm-up), so later-listed or data-gapped legs simply stay silent. The mechanism is new to the corpus. Prior premium runs faded the basis alone at 15m with ~40-minute holds (LINK/ETH/SOL basis fades: 2,000-6,800 trades, liquidated), bought BTC/ETH spot on discount alone (PF 0.37), or traded premium-confirmed momentum/continuation. None conditioned on a sigma-scale price dump and a simultaneous discount, and none used a 12h clock. Stripped mechanism sentence (lesson 138): 'buy a perp after an extreme 24h drop only when its perp-to-index premium shows the drop was derivative-led, hold a fixed half-day'.

Iteration 4 fixes all three auditor findings. (1) Premium measure: the perp/BINANCE_SPOT 4H close proxy and all 10 spot reference legs are removed. Each leg now reads its own hourly Binance premiumIndexKlines series. The pipeline only loaded `premium_index` for the primary symbol, so I added a per-leg loader following the existing OI/taker/top-trader/funding sibling pattern: supp_collector.load_premium_index_series, plus subprocess_entry._add_per_leg_premium_index, which is gated on the strategy quoting……Show moreShow less

Iteration 4 fixes all three auditor findings. (1) Premium measure: the perp/BINANCE_SPOT 4H close proxy and all 10 spot reference legs are removed. Each leg now reads its own hourly Binance premiumIndexKlines series. The pipeline only loaded `premium_index` for the primary symbol, so I added a per-leg loader following the existing OI/taker/top-trader/funding sibling pattern: supp_collector.load_premium_index_series, plus subprocess_entry._add_per_leg_premium_index, which is gated on the strategy quoting 'premium_index_by_instrument', and the Layer-2 synthetic mapping in strategy_verifier._PER_LEG_SUPP, with a unit test. Each hourly value is stamped at its kline CLOSE, so there is no look-ahead. A leg without premium data stays silent; there is no proxy fallback. (2) Exit: the 5% reduce-only STOP_MARKET is deleted, and the only exit is the fixed 12h clock, as pre-registered. (3) Sizing: leg_fraction is restored to 0.12 with max_concurrent 5 (60% gross), and the old clamps and _param_bounds are removed. The library refinements (OI-flush filter, breadth cap, inverse-vol sizing) are NOT added, because the hypothesis does not call for them. They are candidate A/B tests for a later iteration. Local real-data check (2023-06-01 to 2026-09-20, supp_spec path with the new per-leg key): 142 trades, PF 2.70, win rate 0.67, avg_trade_return_pct 1.75%, total return +31.2%, max DD 1.9%, Sharpe 1.47. These match the pre-study's PF 2.3 and win rate 0.67. Per-trade return x trades x 12% (about +29.8%) reconciles with the total return. The avg trade clears the 0.15% floor even with a doubled 0.20% round-trip cost. Performance: per-bar cost is O(legs x log n), with the premium arrays built once; this run has 10 instrument feeds against the previous 20. The local run took 300s on a host at load 66/18 cores.

Failure code: other (conflict with risk policy). This is not fee_edge and not overfit. The edge is real, but it needs to hold through the drop and cannot live with the mandatory stop of at most 5% or the 40% limit on correlated exposure. The version with a stop was measured in iteration 3 and fell to Sharpe 0.46. Promoting again would burn a third risk-review slot on the same rule. Note for the operator: if fixed-clock event books with bounded per-leg loss (about 2.6% of equity per leg) get an exemption that allows a time stop in place of a price stop, revive this hypothesis_id with the iteration-4 code and max_concurrent=3. Note for the Research Lead: do not re-propose this with a tight price stop.

Implementation

Long-only event panel on 10 Binance USD-M alt perps (ETH, BCH, XRP, LTC, TRX, LINK, ADA, BNB, ATOM, DOGE) on 4H bars. On each aligned 4H close, a leg triggers when its 24h log-return z-score (vs its trailing 1080 4H samples) is <= -2.0 AND its OWN latest hourly Binance premium index (premiumIndexKlines close, via the per-leg supp key premium_index_by_instrument, read point-in-time and no more than 2h stale) has a z-score (vs trailing 1080 4H-sampled premiums) <= -1.5. It also needs more than 360 samples of both series, no open position, and 24h since the leg's last entry. Entry is a market BUY at the signal close, which is the next 4H open. The only exit is a fixed 12h clock from the position's open: no stop, no target. Sizing is 12% of equity x leverage (1.0) per leg, at most 5 concurrent legs (60% gross), with simultaneous triggers ranked deepest first by price_z + prem_z.

Verification Results

Get the infra change reviewed and committed, with its unit test, before the full backtest and optimization run. That way every later stage and the paper node see the same feed.

The per-leg premium feed depends on UNCOMMITTED factory changes in the working tree: src/backtesting/subprocess_entry.py (_add_per_leg_premium_index), src/backtesting/supp_collector.py (load_premium_index_series), src/verification/strategy_verifier.py (the _PER_LEG_SUPP entry), and the untracked tests/test_supp_collector_premium_per_leg.py. I reviewed the loader. It is point-in-time correct: each hourly value is stamped at the kline CLOSE, (ts//1h + 1)*1h, and only 1m-1h rows are bucketed. It is gated on the strategy quoting the key. If the working tree is reset or deployed from git without these changes, supplementary_data has no premium_index_by_instrument, every leg stays silent, and the backtest returns no_trades.

Keep this as pre-registered. The analyst should read the worst 5-concurrent drawdown in the full backtest.

Per-trade risk (L148). The hypothesis pre-registers no stop and a fixed 12h clock, and the code implements exactly that. Gross is hard-capped at 5 x 12% x 1.0 leverage = 60% of equity. There is no dynamic scalar, so worst-case gross is 60% <= venue max leverage. There is a flat state, so this is not a compounding blowup. However, the loss per leg on a -19.4% event is about 2.3% of equity, above the 1-2% guideline. The sandbox's largest loss was -$2.68k.

Acceptable on liquid Binance 4H perps. If waiting_on_extra_legs is non-trivial, also run the ETH time exit from a per-bar hook.

Leg statistics update only inside calculate_signal, which runs behind the base cross-leg alignment barrier. If any one of the 10 legs misses a 4H bar, NO leg updates its price or premium z-stats for that timestamp, and entries are skipped. Exits are protected for the extra legs via on_extra_bar. The primary leg (ETH) exits only in the panel pass, so a missing bar on another leg delays the ETH 12h exit by one bar.

Sandbox numbers are strong: PF 2.77 on 287 trades, win rate 0.69, avg trade 2.30% of notional, and all three vol terciles positive. They are close to the pre-study (PF 2.3, WR 0.67). What to check in the full backtest: (1) the one-year concentration, since 2022 carried the most in the pre-study; (2) the worst correlated cluster, where 5 legs x 12% hit in one market-wide flush. There is no stop, and the -19.4% worst event x 12% is about 2.3% of equity per leg, or about 11% if all 5 legs gap together; (3) results excluding the hand-picked list, i.e. the 13-name superset, as an out-of-universe robustness read.

Backtest Review

Sharpe
1.47
Total return
108.57%
Max drawdown
5.21%
Trades
287
Win rate
69.3%
Profit factor
2.82

The numbers are viable and match the pre-study: 287 trades, Sharpe 1.47 (CI 0.79 to 2.17), PF 2.82, win rate 0.69, and avg trade 2.30% against the 0.15% USD-M floor, about 15x. Max DD is 5.2%, every year 2020-2026 is positive, all three vol terciles are positive, and end_unrealized is 0.

It is not beta (L170). Beta is 0.001, correlation to the benchmark is 0.01, and exposure is 7%. The trades fit the mechanism: long only, with 12h holds.

QA's edge concern is only partly borne out. 2022 is the largest year at 23.1% of about 79% summed, which is not over 60%. The clustered-flush risk is real, though: -6.6% on 2022-01-22 and -3.4% on 2021-05-23.

This code is iteration 1's configuration again: no stop, should_exit False, 5 x 12% = 60% correlated gross. The Risk Officer already rejected that exact form on the hard limits mandatory_stop_loss and max_correlated_exposure_pct 40. Optimizing it would reproduce iteration 1's result and end in a third risk_reject.

The only form that meets the limits (5% stop, 8% per leg, 40% gross) was measured in iteration 3 at Sharpe 0.46, with a CI crossing 0 and PF 1.25. The edge depends on not being stopped out: about 1 in 5 trades breached -8% intra-hold. So the deployable version has no demonstrated edge.

The locked final exam was negative on both reads: Sharpe -0.52 and then -2.19 on 5 trades each. Both are underpowered, but they are the only untouched evidence. Programme FDR failed twice (p 0.0051 and 0.044 against a BH cutoff of 0.0015). The holdout has already been read twice.

L126 and L119: this run has already been through OPTIMIZING twice on this mechanism and is now at iteration 4+. A third optimization of an unchanged mechanism cannot change the outcome.

Analysis

The statistical edge holds up. The optimized run made 290 trades with Sharpe 1.17 (confidence interval 0.40 to 1.97), profit factor 2.17, average trade 2.05% of notional (10x the 0.2% floor), max drawdown 5.4%, and a profit in all 7 years.

The robustness checks pass. Deflated Sharpe is 0.970, PBO is 0.123, the result is not overfit, CPCV out-of-sample Sharpe is 0.667, and the holdout Sharpe is 1.47 on 49 trades (consistent with out-of-sample).

The trades match the hypothesis: all long, entered after a sigma-scale dump with the perp at a discount, and closed on a fixed clock. Five of the six pre-registered criteria were met.

The strategy conflicts with the risk limits. It has no stop-loss, but a stop of at most 5% is mandatory. It runs 60% gross exposure against a 40% limit on correlated exposure. Promoting it would repeat the risk rejection from iteration 1.

A version that complies has already failed. 134 of 290 trades (46%) were down more than 5% at some point before the exit (worst -21.5%). The iteration-3 version with a stop fell to Sharpe 0.46 and profit factor 1.25.

The holdout has been read 3 times, so it is no longer a one-shot test. The strategy does not survive the programme-level false-discovery check, and it missed its own pre-registered target of out-of-sample Sharpe 0.9 (actual 0.667).

Code↔hypothesis misalignment found by the semantic auditor — the code does NOT implement the hypothesis. Re-code the strategy to implement the hypothesis EXACTLY (instrument, timeframe, direction, the named edge/mechanic, sizing). Concrete issues: Sizing does not match the title. The hypothesis and pre-study use 12% per leg (5 × 12% = 60% gross). The code defaults leg_fraction to 0.08 and clamps it to [0.04, 0.08], with matching _param_bounds, so 12% can never be reached and gross exposure is capped at 40%. | The……Show moreShow less

Code↔hypothesis misalignment found by the semantic auditor — the code does NOT implement the hypothesis. Re-code the strategy to implement the hypothesis EXACTLY (instrument, timeframe, direction, the named edge/mechanic, sizing). Concrete issues: Sizing does not match the title. The hypothesis and pre-study use 12% per leg (5 × 12% = 60% gross). The code defaults leg_fraction to 0.08 and clamps it to [0.04, 0.08], with matching _param_bounds, so 12% can never be reached and gross exposure is capped at 40%. | The exit rule has changed. The hypothesis says exit on a FIXED 12h clock, and its pre-study stats (worst single event -19.4%, win rate 0.67, PF 2.3, Sharpe 1.32) assume no stop. on_position_opened adds a venue-resting reduce-only STOP_MARKET at 5% below entry on every leg. The rationale admits this changes the payoff. | The premium measure is different. The hypothesis specifies the Binance `premium_index` supplementary key (premiumIndexKlines, perp minus spot index) and the 'latest hourly premium z'. _update_leg never reads that key. It computes premium = perp 4H close / BINANCE_SPOT 4H close - 1 and z-scores that. The result is a 4H perp-vs-Binance-spot close basis, not the hourly index premium the pre-registered study used, so the premium z distribution and the events that trigger differ. ## Library refinements (from the knowledge library; test them, do not assume them) The library supports the auditor on two points. The premium factor should be built from Binance's real Premium Index Klines rather than a perp/spot close proxy, because the two are defined differently. A tight P&L stop fights a non-trend, reversal edge. It also suggests two filters for leverage flushes: falling open interest, and a cap on panel breadth during market-wide cascades. It also suggests inverse-volatility sizing inside the pre-registered 12% leg. 1. [entry] Use the real premium_index key (hourly), not the perp/spot 4H close basis: Replace premium = perp_close/spot_close - 1 with the Binance `premium_index` supplementary series (premiumIndexKlines) for each leg. On every 4H close, read the latest completed hourly premium close at or before bar.ts via supp_as_of / supp_window, never with abs() timestamp matching. Z-score it against that leg's trailing 1080 x 4H samples (sample once per 4H bar, current sample excluded). Keep premium_z_entry = -1.5. If the key only loads for the primary symbol, the premium key must be requested for all 10 legs. Do not fall back to the spot proxy: when a leg has no premium data, that leg stays silent. Drop the BINANCE_SPOT reference legs. — The auditor flagged that the coded premium measure is not the pre-registered one, so the 279-trade backtest (PF 1.25, win rate 0.52) is testing a different trigger from the pre-study (PF 2.3, win rate 0.67). The APFF article says exactly this: the production premium factor uses historical Premium Index Klines, the simulated `Mark/Index - 1` proxy 'is not defined identically, so factor performance observed in the simulated environment cannot be directly treated as evidence for production performance'. It also says 'if auxiliary history is insufficient, the corresponding factor can remain unavailable' rather than being forced. The same doc defines premium mean reversion as a per-symbol time-series z (`-Z_TS(Premium series)`) on 4H data, which matches the hypothesis design. (source: Put Factors Under Continuous Evaluation: Implementing the APFF Multi-Asset Perpetual Strategy on FMZ p.1) 2. [exit] Remove the 5% resting stop; the exit is the fixed 12h clock at 12% per leg: Delete the reduce-only STOP_MARKET at entry x (1 - 0.05). The only exit is a market close at the first bar close >= ts_opened + 12h. Restore leg_fraction = 0.12 with max_concurrent = 5 (gross <= 60%), matching the hypothesis. Bound tail risk with size, not with a stop: the pre-study's worst event of -19.4% x 12% costs about 2.3% of equity. If the risk-limits clamp on stop distance or gross exposure makes 12% impossible, send that back to the Research Lead as a hypothesis change. Do not silently code 8% plus a stop. — The edge is short-horizon reversal after derivative-led forced selling, not trend. Robot Wealth argues a stop lets 'a signal with no predictive power (your P&L) override a signal that does have predictive power'. A stop only fits when P&L predicts future returns, i.e. in trend systems. A 5% stop on alts that have just made a -2 sigma 24h move will often trigger on the continuation wick right before the reversion the strategy is paid for. That fits the lower win rate (0.52 vs 0.67) and PF (1.25 vs 2.3) against the no-stop pre-study. Carver's piece says stops, if used at all, should be set from volatility and time horizon, not a fixed %. A fixed 5% is neither. (source: Stop Losses: Rethinking Conventional Wisdom - Robot Wealth p.1; What is the right way to set stop losses? p.1) 3. [filter] Open-interest flush confirmation: Add a third condition on each leg: the 24h log change in open-interest notional (Binance metrics sum_open_interest_value, read point-in-time) must be <= 0, i.e. OI fell over the same 24h window as the price dump. Keep the price and premium thresholds unchanged. Run this as an A/B against the base rule and adopt it only if it keeps >= 200 trades and raises avg_trade_return_pct. If a leg has no OI data, skip the filter for that leg rather than blocking entries. — The story is that forced leveraged selling is absorbed and then reverts. Amberdata's liquidation guide lists 'Open Interest Dropping, High Liquidation: many traders already got forced out. That may reduce the chance of another cascade in the near term, though it can sometimes also spark opportunistic reversals'. By contrast, 'OI rising' during a dump means positioning is still building and the cascade can extend. The Leverage Purge piece describes the same thing: a leverage flush in which OI fell 30% while price 'overshoots fundamentals… creating dislocation'. APFF uses the same construct as a factor: price direction x Z_TS(24H log change in OI notional). A falling-OI filter should keep the deleveraging dumps and drop the ones still loading up, which may be what is behind the pre-study's -19.4% worst event. (source: Liquidations in Crypto: How to Anticipate Volatile Market Moves p.1; The Leverage Purge: How $8.55B in Liquidations Reset the Market p.1; Put Factors Under Continuous Evaluation: Implementing the APFF Multi-Asset Perpetual Strategy on FMZ p.1) 4. [regime] Breadth cap during market-wide cascades: Count how many of the 10 legs trigger on the same 4H close. If >= 4 trigger together (a market-wide cascade, not a coin-specific perp-led dump), open only the 2 deepest by price_z + prem_z. Open the others only if they trigger again on a later bar after the leg's cooldown. Keep the overall max_concurrent = 5. — Five concurrent longs taken in the same crash are one correlated bet, and the 12h clock cannot get out mid-cascade. Amberdata's October 10, 2025 anatomy shows the cascade ran for 14 hours in waves. The most intense 40 minutes followed hours of slower decline, and alts fell 20-27%, with 65-70% maximum drawdowns in AVAX and DOGE, while BTC fell 7%. The Leverage Purge piece describes the same unwind in phases over weeks ('the cascade unfolded in waves'). Many simultaneous triggers are the signature of that first wave. The coin-specific perp-led dislocations the hypothesis targets should show up as isolated triggers. (source: How $3.21B Vanished in 60 Seconds: October 2025 Crypto Crash Explained Through 7 Charts p.1; The Leverage Purge: How $8.55B in Liquidations Reset the Market p.1) 5. [sizing] Inverse-volatility leg size under the 12% cap, plus a doubled-cost check: Leg notional = equity x leverage x 0.12 x clip(median_RV_panel / RV_leg, 0.5, 1.0). RV_leg is the sample std of that leg's 4H log returns over the trailing 42 bars (7 days), known at the signal close. median_RV_panel is the median of the same statistic across the 10 legs. This can only shrink a leg below 12%, never grow it. Also report avg_trade_return_pct with a second 0.10% round trip subtracted (0.20% total) and require it to stay > 0.15%. — Legs range from BCH/TRX to DOGE/ATOM volatility, and the 12h reversion of a high-vol alt carries much more dollar tail risk than a low-vol one. APFF sizes so that 'higher historical volatility reduces allocation', with a per-symbol cap. It also promotes a factor only if it 'remains positive under doubled costs', because Amberdata shows that during cascades spreads widened about 30x on average, so realized fills around these events cost more than the 0.10% taker baseline. The current 0.62% average trade should clear this. The check guards against the edge depending on the cleanest fills. (source: Put Factors Under Continuous Evaluation: Implementing the APFF Multi-Asset Perpetual Strategy on FMZ p.1; How $3.21B Vanished in 60 Seconds: October 2025 Crypto Crash Explained Through 7 Charts p.1)

Benjamini-Hochberg at q=0.10 over 346 programme candidates keeps 9. A candidate that does not survive here is not distinguishable from the programme's own noise, however good its individual statistics look.

Risk Review

Risk review rejected: - [limit] stop_loss_pct=0.08 exceeds risk_limits max_stop_loss_pct 5% - [limit] critical issue(s) listed by the risk officer: stop_loss_pct=0.08 (8% of entry price) exceeds config/risk_limits.yaml capital.max_stop_loss_pct = 5.0%. pipeline_proces - [waiver rejected] programme_fdr: The candidate p is 0.0442, 29x the Benjamini-Hochberg cutoff of 0.0015 (8 of 341 kept). The programme-deflated Sharpe is 0.68 against an expected max-of-programme of 0.81. The evidence that would need to outweigh……Show moreShow less

Risk review rejected: - [limit] stop_loss_pct=0.08 exceeds risk_limits max_stop_loss_pct 5% - [limit] critical issue(s) listed by the risk officer: stop_loss_pct=0.08 (8% of entry price) exceeds config/risk_limits.yaml capital.max_stop_loss_pct = 5.0%. pipeline_proces - [waiver rejected] programme_fdr: The candidate p is 0.0442, 29x the Benjamini-Hochberg cutoff of 0.0015 (8 of 341 kept). The programme-deflated Sharpe is 0.68 against an expected max-of-programme of 0.81. The evidence that would need to outweigh the FDR miss is weak: OOS Sharpe is 0.50 against the pre-registered 0.9, the holdout has been read twice, the final exam is negative (Sharpe -2.19 on 5 trades, second read), and the sensi - [critical] stop_loss_pct=0.08 (8% of entry price) exceeds config/risk_limits.yaml capital.max_stop_loss_pct = 5.0%. pipeline_processor normalises 0.08 to 8% and blocks promotion on this regardless of verdict. The code's own clamp also allows up to 0.12 (_param_bounds stop_loss_pct [0.04, 0.12]). Both the parameter and the clamp ceiling must be <= 0.05. About 1 in 5 trades end at the -8.1% stop, so a 5% stop changes the payoff distribution materially and needs a fresh backtest and optimization. - [warning] The final exam is negative. On the locked window 2026-04-14..2026-09-29 the strategy made 5 trades with Sharpe -2.19, PF 0.036, win rate 20% and avg trade -0.44%. That is underpowered (5 < 10 trades) and this is the second read (read_no=2), so it is not decisive, but it is the only genuinely untouched evidence and it points the wrong way. - [warning] Missed its own pre-registration: CPCV average OOS Sharpe is 0.50 against a declared min_oos_sharpe of 0.9. In-sample Sharpe is 1.00, below the advisory 1.5. The holdout has been read twice (holdout_exposures=2), so its 1.05 Sharpe / ratio 2.10 is no longer one-shot evidence. - [warning] The sensitivity sweep is effectively unevaluated. 11 of 13 parameters return Sharpe 0.0 at every variation, including the base value. hold_hours returns 0.0 at 12/13/14h while the chosen 18h gives 0.75. The bar_hours 'cliff' is structural, since only 4H data exists. There is no real evidence of parameter stability around price_z_entry, premium_z_entry, stop_loss_pct or z_lookback. - [warning] Correlated-cascade exposure. Up to 5 long legs on highly correlated alt perps trigger on the same bar, and losses cluster. On 2021-05-19..23 about 15 stop-outs of roughly -$700 each occurred in 5 days. On 2022-11-08/09, 9 stops fired in 2 bars. On 2024-08-05, 4 stops fired in 1 bar. The stress windows china_ban_2021 (-1.63%) and yen_carry_aug_2024 (Sharpe -3.27) confirm this. Gross is capped at 40% of BINANCE equity, which equals capital.max_correlated_exposure_pct = 40%, so the cap is at the limit with no headroom. - [warning] Per QA, reported drawdown and return are diluted about 2x by the untraded BINANCE_SPOT reference venue, whose $100k is counted in the $200k starting capital. The true max DD is about 6.3% against the 15% pre-registered limit at 1x leverage. That is still within the limit, but judge every equity-relative number at 2x. - [warning] The deployed cooldown_hours=12 deviates from the pre-registered 'one event per coin per 24h' (the code default is 24). This permits the clustered re-entries during multi-day cascades that the pre-study de-duplicated away. - [info] Live book: 22 members, 0 in breach, no open_items. Several deployed BINANCE members hold ADA/BNB/XRP/ETH directional exposure (e.g. AdaBinanceDualTimeframeMomentumConfluenceLS, EthVolumeConfirmedMomentumLS), and 7 members show replay_verdict 'diverged'. This long-only selloff-buyer adds long beta to the same coins during crashes. Correlation is low on average (benchmark corr 0.15), but it will be high exactly in the stress episodes. - [info] Positives: leverage is 1x (cap 3x); a venue-resting reduce-only stop is placed on fill; time exits run on each leg's own bar; sizing is equity-relative and clamped; avg trade is 1.50% of notional (well above the 0.2% floor); 7/7 positive years; PBO 0.016; impact is 2.6% of gross; capacity is ample for 8% legs on these perps.

Outcome Summary

A mean-reversion edge that has to sit through large intra-hold drawdowns (46% of trades were down more than 5% before exit) cannot survive a mandatory tight price stop, so check risk-limit compatibility (stop and correlated gross) before spending iterations on optimization.

The analyst abandoned it after 4 iterations because it conflicts with the risk policy. The edge needs no stop-loss and 60% correlated gross, but the rules require a stop of at most 5% and cap correlated exposure at 40%. The compliant stopped version (iteration 3) fell to Sharpe 0.46 and PF 1.25, and the Risk Officer also rejected an 8% stop, the programme_fdr waiver, and flagged a negative final exam (Sharpe -2.19 on 5 trades).

Buy one of 10 Binance USD-M alt perps on 4H bars when its 24h return is <= -2 sigma and, at the same time, its perp-vs-index premium is <= -1.5 sigma (a selloff led by the derivative, not spot), then exit on a fixed 12h clock: long-only, 12% per leg, at most 5 legs.

The initial backtest made 287 trades with Sharpe 1.47, total return 108.6%, PF 2.82, win rate 0.69, avg trade 2.30% of notional and max DD 5.2%. The optimized run (15h hold) made 290 trades with Sharpe 1.17, PF 2.17, avg trade 2.05% and max DD 5.4%, plus holdout Sharpe 1.47 on 49 trades, PBO 0.123 and deflated Sharpe 0.970, but it failed programme FDR.

Backtest and paper results are hypothetical. Trading involves risk of loss.