ETF Backtest Design: Universe, Ranking, Costs, and Evaluation
Summary
This configuration specifies a long-only ETF ranking and rebalancing experiment. It defines a broad 100-ETF universe spanning equity, bond, commodity, and regional exposures, with monthly decisions and next-open execution. The main label forecasts a 21-day return; a five-day label uses a weekly schedule so the backtest cadence better matches its horizon. Positions are selected by rank and equally weighted, with top-k choices tested across several portfolio concentrations.
The setup also declares integer-share trading, starting capital, per-share commissions, and tiered half-spread assumptions because daily price data lack bid-ask quotes. Eight expanding temporal splits, a later holdout period, and model and portfolio allocation sweeps establish the evaluation framework. The configuration explicitly flags that the ETF universe was selected backward-looking, creating survivorship bias, and that the liquidity threshold is not inflation-adjusted. Spread tiers rely on industry knowledge, so cost results depend on assumptions rather than observed quotes. It is an experimental specification, not evidence that the configured strategy is profitable.
Key ideas
- The universe is a diversified set of ETFs, but its backward-looking selection creates survivorship bias.
- Monthly ranking and next-open execution define the default decision and trading schedule.
- A weekly cadence is assigned to the shorter return label to align trading frequency more closely with its horizon.
- Equal-weight top-k portfolios and alternative allocators are compared through staged parameter sweeps.
- Transaction costs include per-share commissions and assumed half-spreads because bid-ask data are unavailable.
Tags
Full text
# setup.yaml
```yaml
strategy_id: etfs
setup_version: v1
universe:
assets:
- ACWI
- ACWX
- AGG
- BIL
- BND
- BNDX
- DBA
- DBC
- DIA
- DVY
- EEM
- EFA
- EMB
- EWA
- EWC
- EWG
- EWH
- EWI
- EWJ
- EWL
- EWN
- EWP
- EWQ
- EWT
- EWU
- EWW
- EWY
- EWZ
- EZA
- FXB
- FXE
- FXI
- FXY
- GLD
- GOVT
- GSG
- HYG
- IAU
- IBB
- IEF
- IEFA
- IEMG
- IJR
- INDA
- ITA
- ITB
- IVE
- IVW
- IWM
- IYR
- JNK
- KRE
- LQD
- MCHI
- MDY
- MTUM
- MUB
- OIH
- PPLT
- QQQ
- QUAL
- RSP
- SCHD
- SDY
- SHY
- SLV
- SMH
- SOXX
- SPY
- THD
- TIP
- TLT
- UNG
- USMV
- USO
- UUP
- VCSH
- VEA
- VGK
- VIG
- VLUE
- VNQ
- VTI
- VTV
- VUG
- VWO
- XBI
- XLB
- XLC
- XLE
- XLF
- XLI
- XLK
- XLP
- XLRE
- XLU
- XLV
- XLY
- XME
- XRT
n_assets: 100
eligibility_rule: point_in_time_adv_10m_annual
eligibility_file: eligibility.csv
eligibility_note: The 100 ETFs were selected backward-looking (survivorship bias). The $10M ADV threshold is not inflation-adjusted. See eligibility.csv for point-in-time (asset, year) membership.
decision:
cadence: monthly_month_end
snapshot: close
execution_delay: next_bar_open
# Per-label decision cadence. `cadence` above is the default and applies to every label not
# named here. fwd_ret_5d predicts a five-session return; held to the monthly grid it was
# carried for about 21 sessions, so the backtest traded a horizon the label does not describe
# and the two could not be compared on equal terms. Adding a label here changes its backtest
# spec identity, so it produces new registered rows and leaves existing ones untouched.
cadence_by_label:
fwd_ret_5d: weekly_friday_close
# Engine-level execution defaults. Single source of truth — Ch16-19 notebooks
# read these via get_backtest_config(); never declare a local INITIAL_CASH
# or share_type constant. Changing values here invalidates every existing
# backtest_hash for this case study (cash + share_type are spec-hash inputs).
#
# ``allocator_lookback`` is the bars-of-underlying-price window applied
# uniformly to every moment-based allocator (inverse_vol, risk_parity, hrp,
# mvo_ledoit_wolf). One value per CS keeps allocators comparable — MVO is
# not granted a longer covariance window than IV by default. Bars are
# counted on the price DataFrame the allocator consumes, not the rebalance
# cadence: ETFs has daily Yahoo OHLCV → 63 ≈ 3 months. Conformal-prediction
# allocators have separate calibration windows handled in ``conformal.py``.
execution:
initial_cash: 100_000 # IBKR retail-cohort default; matches per-share fee tier
share_type: integer # ETFs trade in whole shares
allocator_lookback: 63 # 3 months of daily Yahoo OHLCV
mapping:
class: long_only_rank_and_rebalance
position_state_space: long_only
entry_logic: rank_selection_top_n
sizing: equal_weight
costs:
class: material
# Per-share commission plus tier-based half-spread slippage. Yahoo daily
# OHLCV does not carry bid/ask, so spreads are tier-assigned from industry
# knowledge: penny half-spread for the most liquid mega-ETFs and 2c
# default for the rest. Integer share sizing.
model: per_share_plus_spread
per_share: 0.0035 # IBKR Pro Tiered top tier (<=300K shares/mo)
minimum: 0.0
asset_spreads:
SPY: 0.005
QQQ: 0.005
IWM: 0.005
EFA: 0.005
EEM: 0.005
DIA: 0.005
VTI: 0.005
XLK: 0.01
XLF: 0.01
XLV: 0.01
XLE: 0.01
XLY: 0.01
XLI: 0.01
XLP: 0.01
XLU: 0.01
XLB: 0.01
XLRE: 0.01
XLC: 0.01
default_half_spread_usd: 0.02 # 2c for sector/regional/thematic ETFs
spread_convention: half_spread
source: Industry-knowledge tiers; Yahoo daily OHLCV has no bid/ask.
labels:
primary: fwd_ret_21d
buffer: 21D
variants:
- fwd_ret_5d
variant_buffers:
fwd_ret_5d: 5D
# Vectorized-backtest thinning step per label: number of schedule slots
# to advance per trade so holding periods don't overlap. Authored from
# (schedule cadence, label horizon); add an entry here for any new label.
rebalance_step:
fwd_ret_21d: 1 # monthly_month_end schedule, 21d horizon <= 30d gap
fwd_ret_5d: 1 # monthly_month_end schedule, 5d horizon <= 30d gap
backtest:
rebalance:
# A rebalance is skipped when the per-asset weight change is below
# min_weight_change AND the resulting trade notional is below
# min_trade_value. The benchmark profile disables thresholds so that
# full-universe equal-weight (1/N per asset) rebalances at all.
default:
min_weight_change: 0.005
min_trade_value: 100.0
benchmark:
min_weight_change: 0.0
min_trade_value: 0.0
sweep:
# Iteration controls per stage. ``signal: 0`` means "all predictions";
# downstream stages take the top-N from the upstream stage's rank-1.
# Notebooks read these via get_top_n_predictions(case_study, stage).
top_n_predictions:
signal: 0 # all signal predictions (eq-weight baseline)
allocation: 10 # top-10 model configs by equal-weight baseline Sharpe
cost_sensitivity: 1 # top-1 of {signal+allocation} per label
risk_overlay: 1 # top-1 of {signal+allocation} per label
# Each advancing config contributes its single best checkpoint by baseline
# Sharpe, so ten distinct (family, config_name) fill the ten allocation slots
# and no model occupies two of them. Only the checkpoint is fixed here: the
# allocation stage re-sweeps the whole top_k_grid for every advancing config
# (15_portfolio_management.py:122-168), including mappings that lost the
# baseline stage, so a config advances with its checkpoint chosen and its
# top_k still open.
checkpoints_per_config: 1
# Skip MVO/HRP when allocator runtime is the bottleneck (intraday CSes).
expensive_allocators_skip: false
# All other alternative allocators take the cheap path.
# Ch16 equal-weight baseline selection. Long-only top-k. Percentile
# and quantile axes are commented out; uncomment to add them for any
# label and the loader will synthesize the schemes automatically.
# k is a concentration choice and only means anything against the tradeable
# cross-section, which here is 99 names at the median decision date (p10 82).
# 5 / 10 / 20 is 5% / 10% / 20% of that: a genuinely concentrated low end and a
# diversified high end that still selects. The grid was [10, 20] and had no 5%
# end, which is where every comparable case study starts.
top_k_grid:
fwd_ret_21d: [5, 10, 20]
fwd_ret_5d: [5, 10, 20]
# percentile_grid:
# fwd_ret_21d: [80, 90, 95]
# quantile_grid:
# fwd_ret_21d: [5, 10]
# Ch17 portfolio: reuses top_k_grid above and sweeps over allocators.
# Moment-based allocators (IV/RP/HRP/MVO_LW) all use the CS-level
# ``execution.allocator_lookback`` — no per-allocator vol_window / lookback
# fields. No max-weight cap: each allocator expresses its natural
# concentration profile so the cross-method comparison is honest. MVO uses
# Ledoit-Wolf shrinkage as its covariance regularizer; IV/RP/HRP have
# method-inherent diversification; score_weighted is reported unconstrained.
# Equal weight is the Ch16 baseline and is not repeated here.
allocators:
- {name: score_weighted, method: score_weighted}
- {name: inverse_vol, method: inverse_vol}
- {name: risk_parity, method: risk_parity}
- {name: mvo_ledoit_wolf, method: mvo_ledoit_wolf}
- {name: hrp, method: hrp}
- {name: conformal_weighted, method: conformal_weighted}
# Ch18 cost sensitivity (bps regime).
cost_grid_bps: [0, 1, 2, 3, 5, 7, 10, 15, 20, 30, 50]
# Companion per-share cost regime: dollars per share half-spread. Swept
# alongside cost_grid_bps for CSes whose declared cost model is
# per_share_plus_spread. Values: 0¢, 0.5¢, 1¢, 2.5¢, 5¢, 10¢.
cost_grid_half_spread_usd: [0.0, 0.005, 0.01, 0.025, 0.05, 0.10]
# Ch19 risk overlays.
risk_controls:
position:
- {name: stop_loss_3pct, type: stop_loss, threshold: 0.03}
- {name: stop_loss_5pct, type: stop_loss, threshold: 0.05}
- {name: stop_loss_10pct, type: stop_loss, threshold: 0.10}
- {name: stop_loss_15pct, type: stop_loss, threshold: 0.15}
- {name: trailing_1pct, type: trailing_stop, threshold: 0.01}
- {name: trailing_2pct, type: trailing_stop, threshold: 0.02}
- {name: trailing_3pct, type: trailing_stop, threshold: 0.03}
- {name: trailing_5pct, type: trailing_stop, threshold: 0.05}
- {name: trailing_10pct, type: trailing_stop, threshold: 0.10}
- {name: trailing_15pct, type: trailing_stop, threshold: 0.15}
- {name: trailing_20pct, type: trailing_stop, threshold: 0.20}
- {name: time_exit_10, type: time_exit, bars: 10}
- {name: time_exit_20, type: time_exit, bars: 20}
- {name: time_exit_40, type: time_exit, bars: 40}
evaluation:
n_splits: 8
train_size: 10Y
val_size: 1Y
holdout_start: '2024-01-01'
holdout_end: '2025-12-31'
calendar: NYSE
periods_per_year: 252 # NYSE 5d/wk
modeling:
gbm:
libraries: [lightgbm]
preset: default
device: cpu
# LightGBM's own CPU default. 63 is the GPU default and was carried over with the
# device when these runs moved off the GPU, so every CPU fit was quantizing the design
# matrix into a quarter of the bins the library would have used. Coarser bins are
# faster and lose split points; the reader running this on a CPU gets what the
# documentation describes.
max_bin: 255
latent_factors:
persistent_entities: true
macro_context:
source: alfred_initial_release
policy: alfred_initial_release_close_lagged
version: v1
series:
- dgs1
- dgs2
- dgs3
- dgs5
- dgs7
- dgs10
- dgs20
- dgs30
- vixcls
- YIELD_CURVE_SLOPE
- YIELD_CURVE_5_10
availability_lag_days: 1
alignment: backward_asof
model_kwargs:
ipca:
# 100 -> 1000. The eight production folds converge at 53, 59, 69, 77, 85,
# 97, 107 and 110 alternations, so the old cap fell *inside* its own
# convergence distribution: folds 2 and 3 reported exactly 100 and were
# refused by the convergence guard, while fold 1 cleared by three. Which
# folds published was decided by where the cap happened to land. The
# legacy runner never checked convergence at this scale, so the
# truncation registered results and said nothing. 1000 is chosen against
# that measured maximum rather than inherited: it clears every fold by
# roughly nine times, and it still bounds a genuinely stuck fold to about
# 13 minutes at ~0.76s per alternation, where the library default of
# 10,000 would allow two hours. Folds that settle early pay nothing for
# the headroom, because the alternation stops when it settles.
max_iter: 1000
tol: 0.00001
factor_ridge: 0.01
gamma_ridge: 0.01
sdf:
checkpoint_epochs: [256, 512, 768, 1024] # conditional-relative; publishes global 256..1280
beta_checkpoint_epochs: [256]
beta_default_checkpoint: 256
causal:
treatment: skip_recent_6_1
# Bars the treatment's own construction window spans, which is what the placebo block has
# to cover: permuting skip_recent_6_1 in blocks shorter than this destroys the serial
# dependence the refutation exists to preserve, and the resulting p-value reads like a
# refutation without being one. Declared here rather than inferred, because guessing which
# element of a window list a column was built from puts a wrong number behind a right-looking
# one. Derived from the construction, not chosen:
#
# `close.shift(21) / close.shift(6 * 21) - 1` in 03_financial_features: the six-month
# return that stops a month short. The oldest price it reads is 126 sessions back.
treatment_window: 126
confounders: [vol_21d, vol_126d, regime, yield_curve_slope]
method: walk_forward_dml
# The feature-specification register. One row per family: what it reads, how far
# back, with what delay, and whether it ranks assets against each other (signal) or
# describes the environment a ranking is formed in (state). `03_financial_features`
# builds the matrix from `windows` and audits it against `lookback`; nothing outside
# the notebook could state what this case study's feature set is until it lived here.
features:
windows:
momentum: [5, 10, 21, 42, 63, 126, 189, 252]
volatility: [21, 63, 126, 252]
skip_recent: 21
drawdown: [63, 126]
volume: [21, 63]
ranked: [ret_126d, sharpe_126d, vol_63d]
regime_threshold: 0.005 # 10y-2y spread separating the two curve regimes
# Periods for the ml4t.engineer oscillator, trend and regime calls in C.2, and the
# windows C.3 and C.4 read. Declared here rather than typed into the notebook so the
# register, the warmup audit and the timing figure read one source for every window.
oscillators:
rsi: [7, 14]
macd_fast: 12
macd_slow: 26
adx: 14
cci: [14, 21]
stochastic: 14
aroon: 25
natr: 14
choppiness: 14
hurst: 100
sma: [50, 200]
ema: 26
bollinger: 20
state:
obv_zscore: 63 # on-balance-volume z-score
positive_share: 63 # share of up sessions
extremes: 252 # distance from the 52-week high and low
correlation: 63 # SPY-TLT rolling correlation
curve_zscore: 252 # yield-curve slope z-score, over trading sessions
families:
- name: momentum
pattern: ret_*d|skip_recent_*|mom_accel_*
role: signal
hypothesis: Relative performance persists over a one-month hold
inputs: adjusted close
lookback: 252
lag: 0
frame: time series
representation: simple return, and differences between horizons
failure_mode: reverses at the shortest horizons
- name: risk-adjusted momentum
pattern: sharpe_*d
role: signal
hypothesis: Trend earned with less dispersion repeats more reliably
inputs: log return
lookback: 252
lag: 0
frame: time series
representation: window return over its own dispersion
failure_mode: unbounded when dispersion approaches zero
- name: volatility
pattern: vol_[0-9]*d|vol_ratio_short|vol_ratio_medium
role: state
hypothesis: Dispersion sets how far apart the cross-section can spread
inputs: log return
lookback: 252
lag: 0
frame: time series
representation: annualized standard deviation, and ratios of windows
failure_mode: lags a shock by roughly half its window
- name: oscillator and trend
pattern: rsi_*|macd_*|adx_*|cci_*|stoch_*|aroon_*|sma_ratio_*|ema_ratio_*|bb_pctb_*
role: signal
hypothesis: Price against its own recent path separates trend from exhaustion
inputs: OHLC
lookback: 200
lag: 0
frame: time series
representation: bounded oscillator, ratio to a moving average
failure_mode: saturates in a sustained trend
- name: range and drawdown
pattern: natr_*|chop_*|hurst_*|max_dd_*
role: state
hypothesis: Trending and mean-reverting regimes reward different signals
inputs: OHLC
lookback: 126
lag: 0
frame: time series
representation: normalized range, Hurst exponent, share below trailing peak
failure_mode: the exponent needs a long window to settle
- name: volume
pattern: vol_ratio_[0-9]*d|obv_*
role: state
hypothesis: Participation confirms or contradicts a price move
inputs: share volume
lookback: 63
lag: 0
frame: time series
representation: ratio to trailing mean, z-score of on-balance volume
failure_mode: index events put one date orders of magnitude out
- name: extremes and consistency
pattern: pct_positive_*|dist_52w_*
role: signal
hypothesis: Proximity to a 52-week extreme is itself a positioning signal
inputs: adjusted close
lookback: 252
lag: 0
frame: time series
representation: ratio to a rolling extreme, share of positive sessions
failure_mode: piles up at one during a sustained trend
- name: cross-sectional position
pattern: '*_rank'
role: signal
hypothesis: Only relative standing is tradable in a cross-sectional strategy
inputs: ret_126d, sharpe_126d, vol_63d
lookback: 126
lag: 0
frame: cross-section
representation: percentile within the decision date
failure_mode: discards the level the state families carry instead
- name: cross-asset regime
pattern: corr_spy_tlt_*
role: state
hypothesis: Rotation stops working when equities and bonds stop diversifying
inputs: SPY and TLT log returns
lookback: 63
lag: 0
frame: cross-asset
representation: rolling correlation
failure_mode: one equity-bond pair stands in for the whole universe
- name: yield curve
pattern: regime|yield_curve_*
role: state
hypothesis: The slope of the curve separates the macro regimes rotation depends on
inputs: 10y and 2y constant-maturity yields
lookback: 252
lag: 1
frame: macro
representation: level, z-score, regime indicator
failure_mode: >-
the revised Treasury history, not the initial release, so the lag is right for
publication timing but not for revisions
# What `04_model_based_features` decides, declared here for the same reason the
# feature windows above are: an estimation window is part of a fitted feature's
# definition, so it belongs where the definition lives rather than inside the
# notebook that runs it. Every count is in trading sessions.
#
# A fitted feature is bounded by this schedule and not by a cross-validation
# fold. The parameters behind a value are estimated from sessions strictly
# before it, refreshed on the cadence below, and the same value is produced
# whichever fold later selects the row - so the artifact carries no fold column.
model_based:
hmm:
# Calm and stressed. Two states is what the transition matrix and the
# regime means below are read as, and section D reports them by name.
n_states: 2
# Sessions of market history spent before the first regime model is fitted.
# Three years. A two-state chain has to see both states switch several times
# before its transition matrix means anything, and the market series is one
# series rather than a panel, so this is paid once for the whole notebook.
# The folds roll a ten-year training window back one year at a time, so the
# oldest of them begins on the second session of the panel and no burn-in
# fits before it: this one costs the earliest folds training rows. What it
# must not touch is a scored session, and it does not - it ends 2009-02-04,
# 1,736 sessions before the earliest validation session any configured label
# is scored on, 2015-12-28. 04_model_based_features asserts that and prints
# the training rows each fold loses.
burnin: 756
# How often the chain is re-estimated. A quarter. Regime parameters are the
# slowest-moving thing this notebook fits, and each estimate costs
# `n_restarts` expectation-maximization searches.
refit_every: 63
# Expectation-maximization reaches a local optimum, so each fit is repeated
# from this many starting points and the highest training likelihood is kept.
n_restarts: 10
garch:
# Sessions of an ETF's own returns before its variance model is fitted. Two
# years. Below that the persistence term is estimated from too few
# volatility cycles, and it is the term the feature is most sensitive to.
burnin: 504
# Monthly. Faster than the regime model because a variance model tracks a
# level that moves, which is the property the feature exists to report.
refit_every: 21
```Shown in full with attribution under the source's licence. Licence: MIT
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.