This notebook describes an out-of-sample backtest for a selected crypto perpetual funding strategy. It reuses predictions generated from training history that ends before the holdout period, then applies the chosen strategy configuration, including its…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
742 documents
The document describes a feature pipeline that combines equity prices with summaries of listed options implied-volatility surfaces. Its central hypothesis is that disagreement between option-implied volatility and realized share volatility can help rank…
This document turns cross-sectional ETF predictions into simulated trades. It distinguishes ranking quality, measured by information coefficient, from realized strategy performance: a top-k portfolio depends on the relative score values, rebalance schedule,…
The document shows how to turn weekly Commitment of Traders reports into futures positioning features. It explains the trader categories in the financial futures and disaggregated commodity formats, and why participant groups matter when aggregate net…
The document explains how a stochastic discount factor (SDF) estimates a pricing kernel that should price every asset, rather than estimating common return factors. It describes adversarial training: one network proposes the discount factor while another…
This notebook uses Optuna to tune XGBoost, LightGBM, and CatBoost on a time-split firm-characteristics dataset. Each library’s search treats the loss function, either mean squared error or mean absolute error, as a categorical hyperparameter alongside model…
This benchmark compares pandas and Polars on operations found in financial data pipelines, including rolling features, group calculations, window transformations, filtering, joins, lazy scans, memory use, and string handling. It generates shared synthetic…
This notebook explains how to apply four market impact models in a backtest: no impact, linear impact, square-root impact, and a configurable power law. Each model estimates a signed per-share price move based on order direction, quantity, price, and volume.…
This notebook applies TabM, a tabular neural-network approach, to foreign-exchange pair prediction rows without treating the data as a sequence. It uses shared runner infrastructure to fit preprocessing within each training fold, save declared weight…
This document explains how to turn weekly Commitment of Traders reports into futures positioning features. It outlines the report categories for financial futures and physical commodities, describes how net positions reflect different participant roles, and…
This notebook describes a read-only comparison of validation predictions from several model families on a US equities panel. It first checks that each predefined prediction set is complete, then assesses cross-sectional ranking with the information…
This document explains how hidden Markov models infer unobserved market regimes from returns and recent volatility. It first sets two transparent benchmarks: a volatility index threshold for stress and price relative to a long moving average for trend. It…
This notebook implements a univariate forecast of SPY daily closing prices using raw PyTorch, sktime, and Darts. It compares the practical experience of each interface, including implementation effort, installation constraints, and combined fit-and-predict…
This analysis compares predictive, latent-factor, and causal models for a cross-sectional S&P 500 stock strategy using weekly forward returns. Its feature set combines equity momentum and volatility measures with option information such as implied-volatility…
This read-only assessment reconstructs the selected strategy from configured, full-coverage registry results. It follows the progression from an equal-weight baseline through allocation, risk controls, and transaction-cost sensitivity, then reads the holdout…
This notebook assesses whether an LSTM can use the ordering of ETF feature histories to improve on flat-feature linear and gradient-boosting models. It resolves the declared sequence population against current data, checks eligible funds and fund-date rows,…
This notebook uses a synthetic asset panel to demonstrate Instrumented PCA, where factor loadings depend linearly on characteristics observed before returns. Alternating least squares estimates the characteristic-to-loading map and realized factors. Because…
This notebook explores ARIMA as a feature generator for models that may use its output alongside other predictors. It explains how ACF and PACF plots can suggest model orders, how information criteria can compare candidate orders on training data, and why…
This notebook applies a previously selected NASDAQ-100 trading configuration to predictions generated for an untouched holdout window. The strategy, portfolio concentration, allocation method, rebalance schedule, risk overlay, and trading cost assumption are…
This notebook explains why a daily constant-maturity options series is not the return history of a single tradeable contract. Selecting a new near-the-money straddle each day can change the strike, expiration, or both. In particular, moving to a later…
This reference describes academic factor-return datasets from the Fama-French library and AQR for use in strategy analysis and factor modeling. Fama-French offerings include market, size, value, profitability, investment, and momentum series, alongside…
This case study describes producing an out-of-sample prediction set for an already selected S&P 500 options model. The holdout configuration is fixed using validation results, then fitted again on data ending before the holdout window. A label buffer…
This case study compares predictive models for monthly cross-asset rotation across ETFs spanning equities, fixed income, commodities, currencies, and real estate. Its central lesson is that information coefficient (IC) and trading performance can rank models…
This notebook explains how double machine learning (DML) estimates whether FX momentum affects future returns after adjusting for configured confounders. It distinguishes this intervention question from prediction: predictive models are compared by…