Forward-Gated Replacement of Crypto Forecasting Models
Summary
The paper introduces a deployment policy for forecasting systems that retrain regularly: keep the incumbent model serving while a warm-refitted challenger runs off the serving path, then compare both on the same subsequent week of delayed labels. A challenger replaces the incumbent only if it clears a fixed paired advantage in negative log likelihood. This aims to avoid automatic calendar-based swaps when a new fit may be worse than a model that has continued learning.
In historical replay across two separate Binance episodes, the policy reduced negative log likelihood relative to calendar replacement, automatic promotion, and continuous maintenance. It promoted fewer challengers and reduced deployed model changes. The authors report directional consistency across seeds, trial budgets, promotion margins, an earlier asset panel, and a supervised objective. These are replay results on specified crypto perpetual-futures contracts; the supplied account does not establish performance in live deployment or other markets.
Key ideas
- The policy evaluates a warm-refitted challenger alongside the maintained incumbent before replacing it.
- Both models are assessed on the same next week of delayed labels.
- Promotion requires a fixed paired improvement in negative log likelihood.
- Historical replay reports better forecast loss and fewer deployed model changes than the listed comparison policies.
- Reported robustness checks cover seeds, promotion settings, asset panels, and a different supervised objective.
Tags
Full text
# Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets # Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets Production forecasting systems retrain models regularly, but a retrained candidate does not necessarily outperform a continuously maintained incumbent that has continued to learn. We introduce Shadow Before Swap (SBS), a deployment policy that warm-refits a challenger off the serving path, evaluates it against the maintained incumbent on the same next week of delayed labels, and promotes it only after a fixed paired negative-log-likelihood (NLL) advantage. In historical replay over two nonoverlapping Binance episodes spanning 48 UTC weeks, three seeds, eight underlyings, and two perpetual-futures contract types, SBS reduces NLL by 0.1472% relative to calendar replacement, 0.0755% relative to schedule-matched automatic promotion, and 0.0428% relative to continuous maintenance. The corresponding episode-stratified four-week block intervals are 0.1139%-0.1754%, 0.0521%-0.0980%, and 0.0301%-0.0554%, respectively. SBS promotes 114 of 528 challengers, reducing deployed model changes by 78.4% while improving the serving trajectory. The effect remains directionally consistent across seeds, trial budgets, promotion margins, an earlier 20-asset panel, and a topology-matched supervised objective. SBS thus provides a practical deployment policy that improves probabilistic forecasts while limiting consequential model-state transitions.
Shown in full with attribution under the source's licence. Licence: abstract CC0
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.