This notebook surveys supervised-learning labels using ETF price data. It covers fixed-horizon forward returns for regression or direction classification, time-series rolling percentiles, cross-sectional percentile labels, triple-barrier labels with fixed or…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
742 documents
This notebook aggregates registry results for OLS, Ridge, Lasso, and ElasticNet across nine case studies. It selects complete validation results at each study's primary label, compares mean daily cross-sectional rank information coefficients and HAC…
This notebook evaluates ETF features one at a time by calculating a date-by-date rank information coefficient between feature values and subsequent monthly returns. It screens financial and model-derived features over walk-forward validation dates, then…
This notebook demonstrates position-level exits and portfolio-level controls through constructed examples. Static rules include stop losses, profit targets and time exits; dynamic rules include trailing stops that follow prior highs, tightening trails, and…
This notebook evaluates candidate features for ranking US equities by subsequent returns. It computes daily cross-sectional information coefficients, summarizes their average with uncertainty estimates that account for serial dependence, and adjusts…
This notebook shows how to align macroeconomic observations with the dates traders could actually have known them. It distinguishes the period a value measures from its publication date, estimates release dates from period length and agency lag schedules,…
This notebook explains PCA as a baseline latent-factor model for forecasting ETF returns. It decomposes a training panel of forward returns into leading directions and estimates each fund’s exposure and factor premia. The estimator ignores the available…
This guide describes a daily US equities dataset from NASDAQ Data Link’s Wiki Prices, with adjusted OHLCV data and coverage from 1962 through March 2018. It explains how to obtain the archive with an API key, load it, filter by symbol or date, inspect its…
This utility measures the absolute relative change in an entered options straddle's premium over specified session horizons. It builds a valid straddle premium from paired call and put quotes with positive bids and asks above bids, then follows the exact…
This notebook presents a deterministic method for checking whether backtest and live trading pipelines behave alike. It compares successive stages: features computed from the same bars, predictions from those features, signals given the same position state,…
This analysis compares predictive models for ranking NASDAQ-100 stocks by their next 15-minute return using intraday microstructure features such as spreads, depth imbalance, signed volume, and volatility. It emphasizes selecting comparable prediction sets…
This notebook compares sklearn HistGradientBoosting, XGBoost, LightGBM, and CatBoost for predicting forward ETF returns. It measures cross-sectional information coefficient, training time, and memory use across model-complexity presets, with GPU runs…
This notebook tests whether standard portfolio allocation can improve an every-bar NASDAQ-100 trading strategy that is already burdened by transaction costs. It selects predictions using validation performance on the declared cost-feasible universe, then…
This notebook explains how to design a search tool for a forecasting agent so evidence has a consistent structure, a traceable origin, and an auditable path into the model. A shared client protocol returns typed results across providers and includes a cutoff…
This notebook studies how position sizing affects FX backtests after the model has already selected which currency pairs to trade. It preserves the winning baseline’s predictions, signal mapping, costs, and execution settings, then varies allocation rules.…
This utility builds label artifacts for S&P 500 option straddles using the same symbol, strike, and expiration at entry and exit. It aligns feature dates to subsequent market sessions, constructs five- and ten-session exit dates, and joins call and put…
This notebook applies position-level risk controls to leading ETF allocation combinations while keeping each underlying prediction, concentration, and allocator fixed. It compares stop-losses, trailing stops, and time exits with the original strategy,…
This notebook explains how to build a cross-sectional futures feature matrix from three contract tenors per product. It derives carry and curve curvature from exchange-settled prices, while using roll-adjusted prices for return, momentum, and volatility…
This notebook shows why searching across many signals or strategies makes the top observed result look stronger than its underlying predictive value. A simulation uses factors with no true information to illustrate how selecting the largest information…
This notebook explains how to turn a released, cross-sectionally ranked US firm characteristic panel into a keyed feature matrix for a factor study. It groups inputs into declared families, combines characteristics and interactions without fitting…
This US equities feature study explains how to generate features from estimated models without allowing future data into earlier observations. Its estimation schedule uses a history burn-in, fits parameters only on data before each output block, then…
This notebook applies TSMixer to ETF sequences using one globally shared function across funds. Each example contains an individual fund's history and covariates; temporal and feature interactions are learned across the panel, but one fund's observations are…
This notebook describes a stochastic discount factor (SDF) model for pricing the cross-section of US firm returns. Instead of first estimating factors and then applying them, it directly learns firm-month weights under a no-arbitrage condition: discounted…
This notebook queries backtest registries across nine case studies and assembles comparable tables for downstream analysis. It organizes results by asset class and data frequency, then records performance across signal, allocation, cost, and risk stages,…