This notebook demonstrates the Rademacher Anti-Serum protocol as a way to account for selecting a winner from a class of tested strategies. It estimates empirical complexity from candidate performance paths, illustrating how dependence among candidates…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
715 documents
This notebook synthesizes results from nine market case studies into a cumulative strategy-screening funnel. It tests, in order, whether a model has positive information coefficient, whether its selected configuration has positive validation Sharpe, whether…
This notebook assesses whether total value locked can serve as an alternative-data signal for ether returns. TVL aggregates the dollar value of crypto assets deposited in decentralized finance protocols. Because it is a price-valued stock rather than a…
This notebook explains a supervised autoencoder for predicting the direction of future US equity returns across multiple horizons. Its encoder feeds a reconstruction decoder, an auxiliary classifier, and a main classifier. Joint training combines…
This notebook compares pandas and Polars on operations used in financial data pipelines, including rolling calculations, grouped OHLCV summaries, window statistics, filtering, joins, lazy file scans, memory use and string processing. It generates synthetic…
This notebook describes a TabM workflow for foreign-exchange pair models. TabM applies a small neural network to each decision row rather than treating the observations as a sequence. The notebook takes architecture and checkpoint schedules from…
This notebook develops a financial feature matrix for a cross-asset ETF momentum hypothesis: assets with stronger relative performance may continue to outperform over the following month. It combines trailing returns at several horizons, risk-adjusted…
This document presents a cross-market inventory of model-based feature artifacts from nine case studies. It reads parquet schemas rather than loading their rows, excludes identifier columns, counts feature columns, and groups names by tokens associated with…
This document reframes a five-session equity return prediction by sampling daily data on Fridays. The label remains a five-session return, but on the weekly grid it spans about one model step. The notebook compares direct regression, using a fixed lookback…
This document describes refitting the configuration selected by earlier validation stages on all eligible pre-2021 history, then generating predictions for the 2021 holdout. It derives the training interval from the declared evaluation window, label buffer,…
This notebook evaluates gradient-boosted trees on equity option analytics, where features such as implied volatility, skew, term structure, and variance risk premium encode market expectations. It asks whether a nonlinear model can combine those forecasts…
This notebook tests whether gradient-boosted trees improve cross-sectional ranking across currency pairs beyond a penalized linear model. The FX universe contains pairs sharing currencies, so observations are dependent: a move in one currency affects…
This notebook compares two locally run, open-weight embedding models on passages from recent 10-K filings and a fixed set of financial research queries. It builds two-sentence passages, embeds the same documents and queries with each model, and evaluates…
This notebook constructs price-derived features for a broad US equities panel, including momentum, moving averages, and volatility measures. It is designed to rank stocks against one another, using a tradability screen, per-symbol rolling calculations, and…
This notebook compares three ways to orchestrate a four-phase forecasting workflow: direct composition in Python, a role-prompted CrewAI version, and a LangGraph version that delegates to the same specialist classes as the native implementation. It examines…
This notebook checks whether four-hour spot FX data can support a daily cross-sectional strategy that ranks currency pairs using momentum and carry. It tests whether the declared instruments have prices at each decision point, whether the universe represents…
This notebook screens financial and model-based features for their ability to rank stocks by a forward return. It computes daily cross-sectional information coefficients, estimates uncertainty while accounting for serial dependence, adjusts significance for…
This notebook applies principal component analysis to changes in Treasury yields across maturities. Standardizing changes gives each maturity equal influence, and the resulting components are interpreted as level shifts, steepening or flattening, and…
This notebook describes a double machine learning analysis of the effect associated with an FX momentum treatment after adjustment for configured confounders. Flexible nuisance models estimate the outcome and treatment from those confounders; cross-fitting…
This notebook compares three linear forecasters, a Transformer encoder, and two parameter-free forecasts on daily SPY returns. Linear models map a historical window to a multi-day forecast; variants first separate a smoothed component or account for the last…
This notebook configures an NLinear forecasting run for FX pairs. The model uses a fixed consecutive lookback and subtracts the last observed level, directing its fit toward changes over that window. It resolves the lookback, normalization, device, folds,…
This code defines safeguards for reproducible cross-validation, eligibility tracking, and fold-scoped temporal features. It normalizes fold boundaries, compares requested folds with the boundaries used to create temporal artifacts, and rejects incompatible…
This notebook outlines validation-only diagnostics for a 10-session delta-hedged S&P 500 options return label. It organizes financial predictors into implied-volatility-dependent and independent groups, then compares Ridge models using a single volatility…
This notebook recasts prediction of 21-session ETF forward returns as a binary task: positive returns are labeled up, and all others down. It fits L2- and L1-regularized logistic regression using chronological walk-forward folds with a purge gap. Scaling is…