This document describes building a cross-sectional feature matrix for S&P 500 stocks by combining adjusted share-price histories with summarized option implied-volatility surfaces. It organizes features by role, input, lookback, and observability delay.…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
124 documents
This feature-engineering notebook constructs variables that require information beyond one asset’s price history. For futures, it computes annualized roll yield from contemporaneous near and deferred contract levels, plus term-structure slope and curvature…
This notebook builds descriptive market regimes from monthly Federal Reserve economic data: unemployment, the federal funds rate, the Treasury yield curve, and inflation. It explains how to align series with different release frequencies, transform trending…
The notebook diagnoses how a fixed monthly ETF momentum strategy performed across volatility, trend, and yield-curve conditions. Regime labels are designed to be known before the return they describe: volatility and index trend use prior closes and are…
This notebook builds model-based conditional volatility features for S&P 500 shares with a GJR-GARCH(1,1) model, then compares an option-implied volatility spread measured against realized volatility with one measured against a model forecast. Since…
This CME futures study varies gradient boosting tree capacity and loss function while recording predictions at multiple training checkpoints. Tree leaf count controls how finely a model partitions the feature space, while squared, absolute, and Huber losses…
This notebook builds an end-to-end portfolio allocator that places a Temporal Fusion Transformer style variable-selection network ahead of an LSTM. Separate gated residual networks embed individual features, and learned soft weights combine those embeddings…
This setup document defines a weekly S&P 500 equity research pipeline that uses options market features alongside price-based signals. It specifies the eligible universe, decision and execution timing, and distinct rebalance cadences for labels with…
This notebook compares squared error, absolute error, and Huber loss for gradient-boosted models predicting short at-the-money straddle returns. The target has a capped gain and potentially severe losses, so the fitting loss may affect how well predictions…
The notebook treats each window of returns as an empirical distribution and clusters windows according to one-dimensional Wasserstein distance. For equal-sized samples, sorting gives the optimal quantile matching; the distance therefore reflects differences…
This analysis explains how Binance’s perpetual-futures premium index relates to spot prices and how the exchange transforms that premium into periodic funding. The index uses executable impact bid and ask prices relative to an underlying price index,…
This document outlines a two-pass method for extracting source observations used in an S&P 500 options straddle study. First, it identifies call and put contracts meeting a near-the-money candidate screen based on days to expiration, absolute delta,…
The notebook trains a vanilla autoencoder on standardized hourly returns for a group of crypto perpetual markets. Its encoder compresses the cross-asset return vector into a two-dimensional latent representation, and its decoder reconstructs the input. The…
This document explains why a daily constant-maturity option series is not a return series for a position actually held. It selects same-strike, same-expiration call and put legs that pass liquidity, maturity, volatility-estimation, and delta filters, then…
This document presents shared methods for generating fitted-model features without using future observations. For hidden Markov models, it distinguishes filtered state probabilities, based on observations up to the current time, from smoothed probabilities…
This notebook uses double machine learning to estimate whether deviations in perpetual-futures premiums relate to subsequent eight-hour returns, and whether the estimated relationship differs between high- and low-volatility markets. It describes a panel…
This notebook demonstrates deep hedging for a short European call. It simulates geometric Brownian motion paths, uses Black–Scholes delta hedging as a benchmark, and trains a semi-recurrent neural network to choose underlying positions that minimize a…
This document describes a foreign exchange OHLCV dataset covering G10 majors and crosses, with daily and four-hour bars from OANDA. It outlines how the data can be downloaded, loaded, filtered by pair or date range, and explored through coverage summaries…
This notebook evaluates whether option-market measures can rank future returns across stocks. It fits declared linear models on option-derived features, using walk-forward validation, and compares penalty strengths for ridge, lasso, and elastic net. The…
This notebook explains how GARCH turns volatility clustering into a per-session feature. It first uses return plots and an ARCH-LM test to check whether squared returns depend on their own lags. In a GARCH(1,1) model, the response to a new shock and the…
This study asks whether option-market quantities can rank future stock returns. Its features include implied volatility across maturities, put-call skew, term-structure slope, and the variance risk premium. Because many measures are represented in several…
This document develops a daily feature panel for ranking twenty currency pairs at the New York 5 PM close. It distinguishes rankable signals, such as standardized multi-horizon returns and channel position, from market-state measures such as volatility,…
This notebook examines AAPL trade and quote records during the March 16, 2020 market crash to show how market microstructure changes under stress. It filters the tape to regular trading hours, distinguishes trade prints from national best bid and offer…
The notebook explains value at risk as a loss quantile and conditional value at risk, or expected shortfall, as the average loss beyond that threshold. It estimates one-day tail risk for a broad equity ETF using four approaches: empirical historical…