This exploratory analysis profiles a daily panel of exchange-traded funds across nine groups. It checks symbol coverage through time, compares the panel against the classification dictionary, reviews nulls and zero-volume observations, and evaluates…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
243 documents
This notebook describes how to parse IEX DEEP historical feed messages and maintain a limit order book using aggregated price-level updates. It extracts depth changes, best bid and ask quotes, and trade reports, then uses reconstructed snapshots to examine…
This notebook compares three ways to convert signed return predictions into binary positions: a fixed zero cutoff, a trailing percentile of each symbol’s own scores, and a cross-sectional percentile across symbols. It measures signal activation and state…
This NASDAQ-100 study examines how transaction costs affect a frequently rebalanced equity strategy. It separates two possible responses to costs: restricting the traded universe to more liquid, lower-cost names, and reducing rebalance frequency. Keeping…
This deployment demonstration retrains a Ridge model on historical daily data for a cross-section of major FX pairs, fetches current bars through an Interactive Brokers paper session, ranks the pairs, and builds a long basket. It walks through connecting to…
This tutorial turns ETF cross-sectional signals into long-only, equal-weight top-ten portfolios and compares Ridge regression and logistic classification with momentum and equal-weight baselines. Models are fitted on walk-forward training windows, with a…
This notebook trains a Proximal Policy Optimization agent to liquidate a fixed order over a set horizon. It compares the learned pacing policy with TWAP and an Almgren–Chriss schedule, measuring implementation shortfall and examining when each strategy…
This notebook reconstructs the prevailing national best bid and offer (NBBO) for each AAPL trade during the regular session on March 16, 2020. It chronologically combines quote and trade events, carries the latest bid and ask forward, and removes trades…
This tutorial constructs a continuous futures series from individual ES contracts. It identifies the front month using trading volume, with a monotonic roll constraint, and compares volume-based switching with calendar-based timing. It explains why outright…
This notebook examines AAPL trade and quote records during the March 16, 2020 market crash to show how market microstructure changes under stress. It filters the tape to regular trading hours, distinguishes trade prints from national best bid and offer…
This tutorial builds a monthly ETF rotation simulator using a trailing risk-adjusted momentum ranking and a Treasury yield-curve regime filter. At each month-end, the rule ranks ten funds using their recent price history; risk-on months hold the top three…
This notebook demonstrates an operational cycle for an ETF prediction strategy: refresh market data, recompute financial features, retrain a Ridge model, save deployment artifacts, generate current cross-sectional predictions, replay the predictions through…
This notebook develops a framework for translating market activity into modeled trading costs and strategy capacity. It introduces the square-root impact model, identifies volatility and average daily volume as key inputs, and distinguishes quantities…
This document compares four equity market data sources for studying trades, quotes, and limit order books: AlgoSeek TAQ, Databento market by order, NASDAQ ITCH, and IEX HIST. It explains each source’s granularity, coverage, cost or access conditions, storage…
This notebook isolates results for the LEAN engine from an audit comparing retained real-strategy runs with matching ML4T Backtest profiles. It identifies four supported workloads: ETF allocation, crypto perpetual funding, USD-quoted FX allocation, and a US…
This notebook introduces a workflow that connects standardized ETF data loading, a feature registry, and signal diagnostics. It shows how to discover indicator metadata, compute features using defaults or explicit parameters, and store configurations for…
This study builds a minute level measure of order flow pressure from market by order data for NVDA. It signs order additions, cancellations, and fills by side, weights events by size and distance from the midpoint, and smooths the resulting pressure series.…
This document explains how a forecasting agent’s search interface can constrain its evidence and make its activity auditable. A provider-neutral protocol returns typed results with source, publication date, relevance score, and content, while a cutoff filter…
This notebook converts FX pair predictions into baseline trading results. At each decision time it ranks pairs, takes equal-sized long and short sleeves, and runs the selected prediction configurations and checkpoints through an existing backtest engine.…
This notebook compares two simulations of the same monthly ETF momentum strategy. Both use identical target weights, universe, dates, and total trading cost. One computes returns from lagged weights and asset returns; the other processes orders sequentially…
This notebook compares CSV, Parquet, Feather, and HDF5 storage using a deterministic one-million-row OHLCV panel. It measures write time, materialized read time, file size, and reads that select only two columns. The timing policy uses a single write…
This notebook compares DQN, PPO, and A2C in a simulated cryptocurrency trading environment. It calibrates a GARCH(1,1) volatility process on hourly Bitcoin perpetual-futures returns, then simulates paths that preserve volatility clustering while leaving…
This audit compares trading engines replaying the same frozen model targets and content-addressed historical inputs across ETF allocation, futures, crypto perpetual funding, foreign exchange, and US equities. It evaluates execution parity through matching…
This notebook demonstrates a read-only question-answering workflow over an institutional holdings graph built from 13F filings. Instead of generating database queries from natural language, it routes exact supported questions to hand-written Cypher templates…