Skip to content
All strategies

Strategies

Gold Miners vs Gold Beta-Hedged Residual Reversion: fade a >=2-sigma 10-session GDX-vs-GLD idiosyncratic move, dollar/beta-neutral long-short, 15-session time stop (USEQ 1-DAY)

Outcome: Abandoned

GdxGldBetaHedgedResidualReversion

Outcome Summary

The strategy tried to capture partial reversal of gold-miner moves that gold does not explain, trading GDX against a beta-hedged GLD leg on daily US equity bars. Over 4 iterations the code never passed verification: Layer 1.5 kept flagging a hypothesis/config mismatch because the optimization_plan's exit_z baseline was missing or outside its domain. The run was abandoned in that verification loop, so backtesting, optimization, analyst review and risk review were never reached and no performance metrics exist.

Hypothesis

Two-instrument long-short pair on USEQ daily bars, OHLCV only. GDX starts in 2006 and GLD in 2004, so the overlap is about 19 years. GDX (gold miners) behaves like a levered claim on gold (GLD) because of the miners' operating leverage. Each session the strategy computes: (1) a rolling OLS beta of GDX daily log returns on GLD daily log returns, using data up to the previous session; (2) the daily residual; (3) the residual summed over the last resid_window sessions; and (4) a z-score of that sum against the prior 250 sessions. Large miner moves that gold does not explain are expected to partly reverse within weeks, because miner cash flows stay tied to the gold price. Examples are equity risk-off selling, ETF creation/redemption pressure and sector rotation. This design avoids the dead families in the lessons: the trigger is not a calendar date (L179), there is no crypto supplementary-feed gate (L111), and it uses no options (L102). Leverage is not a research parameter. Gross exposure comes from equity through the fixed gross_exposure_pct.

This fills under-represented buckets: USEQ venue (5.4%), long-short direction (16.3%) and pairs scope. It also moves away from BINANCE (62.5%) and BTC. A zero-commission venue and multi-percent moves per trade address the fee_edge graveyard. About 19 years of daily data gives a large sample, and about 5,000 bars keeps the sandbox run short. Daily, 1-hour and 1-minute bar folders for both GDX.USEQ and GLD.USEQ are in the catalog, but the overlapping daily history was not checked. Leverage is left out of optimization_plan because three recent runs died at Layer 1.5 on 'optimization_plan fixes leverage'.

verification_loop: Verification failed (Layer 1.5 — hypothesis/config consistency) [class=hypothesis_mismatch]: - optimization_plan baseline exit_z is missing or outside its domain

Implementation

Continuously sizes a GDX/GLD beta-hedged pair against prior-session residual z. Rebalances beyond a 0.25 full-allocation band, flattens on sign crossings, correlation below 0.5 or |z| above 3.5, and retains the 6% pair stop, 15-session horizon and three-session cooldown. Target gross exposure is capped at 80% of marked equity.

The supplied previous_code already implements the requested continuous mechanism; the smallest correction makes calculate_signal report the prior-session z actually used for trading, aligning subsequent IC analysis with execution timing. Rolling OLS, residual construction, synchronized daily bars, sizing and risk logic are preserved. Evaluation targets are recorded as metadata for pipeline assessment, not claimed as achieved or automatically enforced. The research plan still lists obsolete entry_z/exit_z tunables and does not authorize tuning full_size_z/rebalance_band; it needs a Research Lead revision before meaningful continuous-sizing optimization. Optional taper, jump veto, dynamic half-life and Kalman refinements were not added because they broaden scope or conflict with fixed parameters. No backtest or final-exam data was read. The requested outbox file could not be written because this session permits filesystem reads only; the complete bound JSON is returned here.

Verification Results

Verification failed (Layer 1.5 — hypothesis/config consistency) [class=hypothesis_mismatch]: - optimization_plan baseline exit_z is missing or outside its domain

Optional: note the gap risk, or reduce gross_exposure_pct if the drawdown floor binds.

The 6% pair stop is checked only at the daily close and exits at market on the next processing. A gap across one or more sessions can overshoot it materially; the sandbox's largest loss of -$8958 against an average loss of -$1178 is consistent with this. The hypothesis does not specify a stop at all, so this is not a mismatch. Worst case at gross 0.8x: GDX notional is at most 0.8/1.3 ≈ 0.62x equity, so a 6% stop risks about 3.7% of equity before gap overshoot.

Also return early from _manage() when getattr(self, '_no_entries_before_ns', 0) > clock ns, or set _side only after fills are confirmed.

_manage() skips order submission only under _in_warmup. Under the parity replay seed freeze (_no_entries_before_ns), _submit_entry_instrument silently drops the orders, but _side, _entry_* and _q_* are still set. On the next bar the 'flat with side != 0' branch resets _side and starts a cooldown. Replay state can therefore differ from paper by a 3-session cooldown. Normal backtests are unaffected.

Acceptable as written. Optionally log how often the clamp binds.

The estimated beta is clamped to [0.3, 4.0]. When the clamp binds, the residual uses alpha = my - b*mx with the clamped b, so it is no longer the OLS residual. This only matters in degenerate windows and is documented as a structural guard.

Analysis

Change the mechanism rather than retune parameters. 76 pair events are too few to optimize on, and the holdout would see only about 15. (1) Size the position continuously from z: target GDX notional = -clip(z/2, -1, 1) × G/(1+β), hedged with GLD at the estimated beta. Rebalance only when the target moves by more than 0.25 of full size. Aim for at least 60% time in market and more than 300 rebalance events. (2) Add a veto that sets the target to zero when the 60-day GDX/GLD return correlation is below 0.5, or when……Show moreShow less

Change the mechanism rather than retune parameters. 76 pair events are too few to optimize on, and the holdout would see only about 15. (1) Size the position continuously from z: target GDX notional = -clip(z/2, -1, 1) × G/(1+β), hedged with GLD at the estimated beta. Rebalance only when the target moves by more than 0.25 of full size. Aim for at least 60% time in market and more than 300 rebalance events. (2) Add a veto that sets the target to zero when the 60-day GDX/GLD return correlation is below 0.5, or when |z| is above 3.5 (the residual is trending rather than reverting). (3) Keep beta and z computed from prior sessions only, the 6% pair stop, and gross exposure at or below 0.8 × equity. (4) Set these targets in advance: full-history Sharpe at least 0.6, no single year above 40% of profit, and calm and normal regimes both non-negative. If the IC t-stat is still below 2, abandon the premise as falsified. ## Library refinements (from the knowledge library; test them, do not assume them) The library covers GLD-GDX directly. Hudson & Thames / ArbitrageLab use this exact pair to show stochastic-control pair trading, where the optimal spread weight is linear in the mispricing and is cut back outside a 'stabilization region'. That supports replacing the 2-sigma on/off trade with continuous z-proportional sizing plus a taper at extreme z. Pairs-trading reviews add a half-life-based time stop, a veto on idiosyncratic-news jumps and a dynamic (Kalman) hedge ratio. These all address the current weak result (Sharpe 0.23, PF 1.19, 76 pair events). 1. [sizing] Continuous z-proportional pair sizing instead of 2-sigma on/off: Drop the entry_z/exit_z state machine. Each session, after z is formed from prior-session stats, set target GDX notional = -clip(z/2.0, -1, 1) * G/(1+beta), with G = equity * 0.8, and hedge GLD at +beta times the GDX notional with the opposite sign. Rebalance both legs only when |target_frac - current_frac| > 0.25 of full size, or when the target crosses zero (go flat). Keep the 6% pair stop on GDX-leg entry notional (after a stop, the target is 0 for reentry_cooldown=3 sessions). Tunables: full_size_z in [1.5, 3.0] (default 2.0) and rebalance_band in [0.15, 0.35] (default 0.25). — In the Jurek-Yang and Mudchanatongsuk models (both shown on GLD/GDX), the optimal spread weight h*(t,x) contains the term -k(x-theta)/eta^2, which is linear in the mispricing. The 'optimal portfolio weights are dynamic', so the position scales with divergence rather than switching at one threshold. Our binary 2-sigma entry produced only 76 pair events, too few to optimize or holdout-test. A graded position puts capital to work at |z| 0.5-2 and multiplies the decision count, which is what the analyst's ≥300 rebalance target needs. Fees: 0.25-of-size steps on 0.8x gross cost little per step on USEQ, which has zero commission and only spread/impact costs. The library's equity curves are claims; our backtest decides. (source: Pairs Trading with Stochastic Control and OU process - Hudson & Thames p.1; ou model mudchanatongsuk (ArbitrageLab docs) p.1; ou model jurek (ArbitrageLab docs) p.1) 2. [regime] Stabilization-region taper: shrink the position when divergence is extreme: Multiply the target from refinement 1 by taper(|z|): 1.0 for |z| <= 3.0, linear down to 0 at |z| = 4.0, and 0 above that. Once the taper has cut the position to 0, do not re-open until |z| has fallen back below 3.0. All z values come from prior-session statistics only. — Jurek-Yang's result, as summarised in the article: 'there is a critical level of the mispricing beyond which further divergence ... precipitates a reduction in the allocation'. Outside the stabilization region 'the wealth effect dominates, leading the agent to curb his position'. That is a principled version of the analyst's |z|>3.5 veto. A residual that has already moved more than 3 sigma is more likely a regime break than noise, and that is where the 6% pair stop gets hit. The article illustrates stabilization bounds on GLD-GDX in Fig. 4, but that figure is not stored in the library, so I rely on the text only. (source: Pairs Trading with Stochastic Control and OU process - Hudson & Thames p.1; ou model jurek (ArbitrageLab docs) p.1) 3. [filter] Idiosyncratic-jump veto (news-driven divergence continues): Each session, compute sigma_d = std of daily residuals over the prior beta_lookback=60 sessions. If any single-day residual inside the current resid_window has |resid| > 4.0 * sigma_d, set the target to 0 and keep it at 0 until that day leaves the window. This does not stop an open position by itself; the target simply goes flat at the next rebalance. — Robot Wealth's literature review reports that pair-trading profit 'is related to idiosyncratic news. News affecting one of the stocks in isolation is predictive of a divergence that continues, rather than converges.' Our residual sum S cannot tell ten small drifts (flow pressure, which should revert) from one large one-day gap (a miner-sector shock, M&A or a guidance change, which may not revert). Vetoing windows that are dominated by one jump targets the losing trades directly. The library gives no threshold; 4 sigma is our pre-registered choice. (source: Pairs Trading Literature Review - Robot Wealth p.1) 4. [exit] Half-life-matched horizon: z window and time stop from the OU half-life: Each session, fit AR(1) to the cumulative residual S over the prior 250 sessions (dS_t = a + b*S_{t-1}) and compute HL = -ln(2)/ln(1+b). If b >= 0, treat the spread as non-reverting and set the target to 0. Set the time stop to max_hold = clip(round(2*HL), 5, 30) sessions: if the target has kept the same sign for longer than max_hold, force it to 0 and apply the 3-session cooldown. Replace the fixed max_hold=15 with this rule and log HL so the analyst can see it. — The QuantInsti stat-arb project estimates the half-life from a lagged regression of the spread, then builds its z-score over 'half-life' intervals. Robot Wealth reports that 'return potential decreases significantly with time after divergence', hence the rule to 'puke trades that haven't converged after a period of time'. Our fixed 15-session stop ignores how fast the GDX/GLD residual actually reverts in each era. A b >= 0 check also gives an in-strategy test that the premise still holds, which complements the analyst's IC t-stat test. (source: Kalman Filter Techniques And Statistical Arbitrage In China's Futures Market In Python p.1; Pairs Trading Literature Review - Robot Wealth p.1; Introduction to Pairs Trading (Lecture 42) p.1) 5. [parameter] Kalman-filter hedge ratio instead of 60-day rolling OLS: Estimate (alpha, beta) of r_gdx on r_gld with a 2-state random-walk Kalman filter. Use transition covariance delta/(1-delta)*I with delta=1e-4, and observation variance equal to the residual variance of the first 60 sessions; after that, update it with an EWMA of squared innovations over a 60-session span. Use the PRIOR-step state estimate to form today's residual and hedge, then update. Keep the clamp beta in [0.3, 4.0]. Use the same beta for the GLD hedge leg. — The QuantInsti project replaces fixed-window regression with a Kalman filter for the hedge ratio: 'there's no window length that we need to specify' and it 'weigh[s] recent observations more'. ArbitrageLab's hedge-ratio module notes that OLS hedge ratios depend on which leg is the dependent variable (TLS, Johansen and minimum-half-life are alternatives). Its GLD/GDX example shows OLS with GLD as the dependent variable. Robot Wealth notes the regression slope 'tends to be quite unstable'. Each jump in a 60-day OLS beta leaks gold beta into the residual, which creates fake z extremes. Confidence is low because the library only claims this is better; our backtest decides. (source: Kalman Filter Techniques And Statistical Arbitrage In China's Futures Market In Python p.1; hedge ratios (ArbitrageLab docs) p.1; Pairs Trading Literature Review - Robot Wealth p.1)

Iteration History

Verification failed (Layer 4 — QA review) [class=hypothesis_mismatch]: - [edge_concern] The sandbox produced 328 closed positions with negative net returns and profit factor 0.798. Full-history analysis should assess convergence after spread, impact and borrow costs; these results do not determine this correctness verdict. - [critical] The supplied hypothesis requires fading residual moves of at least 2 sigma. The code instead continuously targets exposure proportional to z, without consulting entry_z. With……Show moreShow less

Verification failed (Layer 4 — QA review) [class=hypothesis_mismatch]: - [edge_concern] The sandbox produced 328 closed positions with negative net returns and profit factor 0.798. Full-history analysis should assess convergence after spread, impact and borrow costs; these results do not determine this correctness verdict. - [critical] The supplied hypothesis requires fading residual moves of at least 2 sigma. The code instead continuously targets exposure proportional to z, without consulting entry_z. With rebalance_band=0.25 and full_size_z=2, a flat book can enter when |z| exceeds approximately 0.5, subject to whole-share sizing. The rationale refers to iteration feedback authorizing this replacement, but that feedback is absent from this task's hypothesis. (line 237) — fix: Implement the supplied >=2-sigma entry gate and the authorized exit structure, or obtain a revised task-bound hypothesis explicitly specifying continuous exposure and its vetoes. - [critical] The rebalance early return checks only the GDX exposure fraction and total gross notional. It never checks GLD's deviation from the current beta-weighted hedge. A change in beta can therefore leave a materially incorrect hedge while the GDX difference remains within the band and gross remains within budget. For example, after beta falls, the existing GLD short can remain oversized indefinitely despite the stated beta-neutral sizing. (line 307) — fix: Compute desired whole-share quantities for both legs before the early return. Require both legs' target deviations and the residual beta exposure to satisfy explicit tolerances before skipping a rebalance. - [warning] The attached optimization plan tunes entry_z and exit_z, but neither parameter affects this implementation. Trials varying those dimensions would measure identical behavior and would not test the declared entry/exit parameterization. (line 53) — fix: Reconcile the implementation and optimization plan with the authorized hypothesis before optimization; do not substitute new parameter permissions implicitly.

Outcome Summary

Every parameter named in the research optimization_plan (here exit_z) must exist in the strategy's parameters with a baseline inside its declared domain, or the run cannot get past Layer 1.5 verification.

After 4 iterations it was abandoned in a verification loop, failing Layer 1.5 (hypothesis/config consistency, class hypothesis_mismatch) because the optimization_plan baseline for exit_z was missing or outside its domain.

A dollar/beta-neutral GDX-vs-GLD long-short pair on USEQ daily bars that fades a >=2-sigma 10-session idiosyncratic residual move of gold miners vs gold (rolling OLS beta, z-scored against the prior 250 sessions), with a 15-session time stop.

No performance data exists: the backtest report is empty and no optimization was run, so there are no return, Sharpe or trade figures.

Backtest Review

Sharpe
0.23
Total return
24.48%
Max drawdown
12.19%
Trades
152
Win rate
51.3%
Profit factor
1.19

Backtest and paper results are hypothetical. Trading involves risk of loss.