This notebook tests Lee-Ready trade classification against aggressor-side labels in Nasdaq order-by-order data. It reconstructs the limit order book from add, modify, cancel, fill, and reset messages, then aligns each trade with the contemporaneous best bid…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
566 documents
This notebook uses TreeSHAP to explain LightGBM return predictions, with global feature importance, individual prediction explanations, dependence plots, and pairwise interaction values. It compares SHAP rankings with permutation importance and tree-based…
This notebook turns financial headlines into stock-level signals and evaluates them against forward returns. It embeds headlines, measures news surprise as semantic distance from a rolling embedding baseline, and combines surprise with sentiment direction to…
This document explains how to evaluate a previously selected US equities strategy on holdout data while keeping its configuration fixed. It derives the correct holdout prediction set from the selected model’s training identity and checkpoint, then applies…
The notebook demonstrates a pipeline for extracting supplier, customer, and competitor relationships from company annual filings and storing them as a knowledge graph. A local language model generates candidate subject–relationship–entity triples from filing…
This notebook measures how trading costs affect strategies already selected through signal, allocation, and risk-overlay stages. It fixes one validation-selected configuration per label before varying costs, so the resulting curves isolate the effect of the…
This document explains why stop-losses, trailing stops, and time exits cannot be evaluated in a case study whose backtest holds weights across a month and observes only the realized monthly forward return. Such rules depend on the price path between entry…
This document describes a strategy assessment process for registered NASDAQ-100 microstructure backtests. It reads existing runs rather than training models or rerunning backtests, traces selected strategy lineages, compares candidates with an equal-weight…
This document presents a staged method for matching company names from alternative data, filings, news, and price sources to securities. It first joins on available identifiers in a deliberate trust order, preserving unmatched records and checking that…
This document describes building a cross-sectional feature matrix for S&P 500 stocks by combining adjusted share-price histories with summarized option implied-volatility surfaces. It organizes features by role, input, lookback, and observability delay.…
This notebook defines monthly total-return labels for a cross-sectional US firm-characteristics study and checks how those outcomes align with the provider’s rows. The panel pairs characteristics from the prior month with the return earned in the month…
This notebook shows how explicit execution costs can erode a hypothetical intraday strategy’s gross returns. It anchors the crossing spread to the median volume-weighted quoted spread across NASDAQ-100 constituents, then builds crossing, worked-order, and…
This notebook explains how a temporal convolutional network (TCN) can forecast returns from ordered financial data. Causal convolutions ensure that an output at a given time depends only on current and earlier inputs. Dilations expand the receptive field…
This document presents an event-driven method for holding a limited number of intraday positions. Predictions are aligned to price bars using the latest available score, subject to an optional freshness limit. Entry signals use a rolling quantile computed…
This case study turns stored model predictions for US stocks into long-short, equal-weight portfolios. It sweeps multiple entry schemes, each selecting the highest-ranked names for long positions and the lowest-ranked names for short positions, to examine…
This live-trading demonstration describes a daily rebalance workflow for a fixed universe of US large-cap stocks. It compares broker-held positions with model targets, converts the differences into an order basket, and routes orders through a risk-control…
This notebook compares sklearn HistGradientBoosting, XGBoost, LightGBM, and CatBoost on an ETF return-prediction task. It measures cross-sectional information coefficient, training time, and process memory across CPU and, where supported, GPU runs, using…
This notebook turns cross-sectional stock predictions into long-short portfolios by ranking stocks at each rebalance, buying the top k and shorting the bottom k with equal capital per position. Equal weighting provides a baseline that isolates differences in…
The notebook previews S&P 100 10-K and 8-K filing data intended for financial knowledge graph construction. It checks schema completeness, unique company and accession keys, form consistency, filing year, and recorded text length. Annual reports supply…
This notebook evaluates model predictions as NASDAQ-100 trading strategies using a shared backtest engine. It first runs a plumbing check with random signals: persistent profits after costs would point to issues such as lookahead, misaligned data, or…
This notebook checks whether a fifteen-minute cross-sectional stock-ranking strategy is feasible before fitting a model or making forecasts. It uses NASDAQ-100 quote data, midpoint returns, and a pre-holdout development window to examine the strategy’s…
The document describes fitting a simple linear sequence model to NASDAQ-100 one-minute features to forecast forward returns at several horizons. Each input contains the preceding hour for one stock, expressed as changes from the latest observation, and the…
The document describes loading ETF market data with optional symbol and date filters, plus a deterministic limit on the symbols returned. Its main analytical point is to match the price series to the quantity being measured: adjusted prices are appropriate…
This notebook builds model-based conditional volatility features for S&P 500 shares with a GJR-GARCH(1,1) model, then compares an option-implied volatility spread measured against realized volatility with one measured against a model forecast. Since…