This notebook constructs price-derived features for a broad US equities panel, including momentum, moving averages, and volatility measures. It is designed to rank stocks against one another, using a tradability screen, per-symbol rolling calculations, and…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
566 documents
This notebook builds minute-level features from NASDAQ-100 quote and trade data to study short-horizon price pressure. It treats normalized order-flow imbalance as the main signal candidate and uses spread, depth, price impact, off-exchange trading,…
This notebook screens financial and model-based features for their ability to rank stocks by a forward return. It computes daily cross-sectional information coefficients, estimates uncertainty while accounting for serial dependence, adjusts significance for…
This notebook compares three linear forecasters, a Transformer encoder, and two parameter-free forecasts on daily SPY returns. Linear models map a historical window to a multi-day forecast; variants first separate a smoothed component or account for the last…
This notebook recasts prediction of 21-session ETF forward returns as a binary task: positive returns are labeled up, and all others down. It fits L2- and L1-regularized logistic regression using chronological walk-forward folds with a purge gap. Scaling is…
This notebook tests whether an LSTM can extract temporal structure from ETF feature histories that flat-feature models may miss. It resolves the declared sequence population against available features, labels, entities, and walk-forward folds before fitting.…
This notebook defines close-to-close forward returns over two trading horizons for a cross-section of ETFs, with each horizon measured from an adjusted close to the close a fixed number of sessions later. It explains why labels are built on complete symbol…
This case study builds minute-level features from NASDAQ-100 quote and trade data to examine whether recent aggressive buying or selling predicts short-horizon price drift. Order-flow imbalance is the proposed signal; spread, book depth, price impact,…
This notebook constructs several sampling schemes from a single day of NASDAQ ITCH trades for an equity: calendar-time, tick, volume, dollar, imbalance, and run bars. It compares their statistical properties, including normality and autocorrelation, and…
This notebook demonstrates two sequential monitors on validation prediction errors from two linear equity model configurations: a two-window mean-shift detector and a monitor for the frequency of bad days. Each detector is calibrated on an initial period and…
This notebook evaluates four signals derived from news text: weighted surprise, average sentiment, sentiment momentum, and article coverage. It uses forward returns prepared by an earlier feature-building step, then calculates a daily cross-sectional…
This notebook applies double machine learning (DML) to estimate the effect of skip-recent momentum on ETF forward returns, a causal question distinct from forecasting returns. It models the outcome and the momentum treatment using declared confounders, then…
The notebook demonstrates tuning LightGBM for ETF return prediction with Optuna’s TPE sampler. It searches tree structure, sampling, and regularization settings, using early stopping to choose the boosting rounds and a custom callback to report…
This notebook turns institutional 13F holdings into a bipartite institution-to-stock graph and derives features for research, including stock co-ownership similarity, ownership breadth and concentration, and changes in reported holdings. It aggregates…
This notebook compares exhaustive grid search with Optuna’s TPE Bayesian sampler for tuning a LightGBM model on an ETF forward-return prediction task. It first gives both methods the same trial budget over a small categorical grid, where exhaustive search…
This notebook studies how to calibrate time, tick, volume, dollar, and imbalance bars using multiple sessions of NVDA market-by-order trade data. It filters trades to regular trading hours, uses the feed’s aggressor-side labels, and examines day-to-day…
This notebook defines forward-return labels for a US equities panel and explains why their construction affects every downstream model and backtest. It specifies adjusted-price return windows in trading sessions, checks that each stock has the required…
This notebook trains TabM neural networks to rank ETFs using the same flat feature table as linear and boosted models. Each ensemble member shares a two-layer backbone but has its own scaling vector and output layer, allowing predictions to be averaged with…
This notebook explains how to decompose ETF returns and risk using CAPM and Fama–French factor regressions. It estimates full-sample exposures with heteroskedasticity and autocorrelation robust standard errors, tracks changing betas with rolling windows, and…
This notebook adapts skip-gram Word2Vec to institutional holdings by treating each 13F portfolio as a sentence, each stock identifier as a token, and position size rank as token order. Nearby positions form the context, so stocks that institutions place in…
This notebook describes fitting PatchTST to one-minute NASDAQ-100 data to predict returns over several forward horizons. The model groups consecutive observations into patches and applies attention across them, reducing the number of items compared while…
This notebook compares ways to allocate capital across US equity positions while holding the model, checkpoint, rebalance dates, and selected stocks fixed. It examines weights based on prediction strength, prediction-interval width, individual stock…
This notebook explains an unconditional signature-based Wasserstein GAN for generating financial time series. It transforms returns into augmented paths, computes truncated path signatures, and trains an LSTM generator driven by Brownian noise to match…
This notebook evaluates stop losses, trailing stops, and fixed-duration exits on selected US equity strategies. A stop loss responds to losses from entry, a trailing stop responds to declines from a position’s peak, and a time exit closes after a set holding…