The document describes a feature pipeline that combines equity prices with summaries of listed options implied-volatility surfaces. Its central hypothesis is that disagreement between option-implied volatility and realized share volatility can help rank…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
566 documents
This document turns cross-sectional ETF predictions into simulated trades. It distinguishes ranking quality, measured by information coefficient, from realized strategy performance: a top-k portfolio depends on the relative score values, rebalance schedule,…
This document describes a daily ETF candidate universe covering equities, fixed income, commodities, and currencies. It outlines a workflow for downloading market data, loading it for analysis, inspecting coverage by symbol and category, and filtering by…
This notebook describes a read-only comparison of validation predictions from several model families on a US equities panel. It first checks that each predefined prediction set is complete, then assesses cross-sectional ranking with the information…
This notebook implements a univariate forecast of SPY daily closing prices using raw PyTorch, sktime, and Darts. It compares the practical experience of each interface, including implementation effort, installation constraints, and combined fit-and-predict…
This notebook explains a family of ETF models that represents returns through shared latent directions and estimates how fund features map to exposures. It distinguishes five approaches: unconditional principal components, instrumented PCA with a linear…
This analysis compares predictive, latent-factor, and causal models for a cross-sectional S&P 500 stock strategy using weekly forward returns. Its feature set combines equity momentum and volatility measures with option information such as implied-volatility…
This notebook describes reconstructing a single-symbol, single-day NASDAQ limit order book from message-by-order ITCH data. It processes add, execute, cancel, delete, and replace messages while maintaining each live order reference and its remaining shares.…
This notebook profiles an annual-report corpus before indexing it for financial research. It highlights four data issues that row counts and null checks do not reveal: the gap between fiscal year-end and public filing, duplicate records when one filing…
This read-only assessment reconstructs the selected strategy from configured, full-coverage registry results. It follows the progression from an equal-weight baseline through allocation, risk controls, and transaction-cost sensitivity, then reads the holdout…
This exploratory analysis explains how to interpret minute bars built from quote and trade data for NASDAQ-100 constituents. It organizes the fields into families covering bid and ask quotes, executions, spreads, volume, trade-price buckets, tick direction,…
This notebook explores ARIMA as a feature generator for models that may use its output alongside other predictors. It explains how ACF and PACF plots can suggest model orders, how information criteria can compare candidate orders on training data, and why…
This notebook applies a previously selected NASDAQ-100 trading configuration to predictions generated for an untouched holdout window. The strategy, portfolio concentration, allocation method, rebalance schedule, risk overlay, and trading cost assumption are…
This reference describes academic factor-return datasets from the Fama-French library and AQR for use in strategy analysis and factor modeling. Fama-French offerings include market, size, value, profitability, investment, and momentum series, alongside…
This document outlines a shared data system for quantitative trading research, cataloging datasets across equities, options, futures, crypto, foreign exchange, factors, macroeconomics, filings, positioning, news, and prediction markets. It describes the…
This notebook surveys supervised-learning labels using ETF price data. It covers fixed-horizon forward returns for regression or direction classification, time-series rolling percentiles, cross-sectional percentile labels, triple-barrier labels with fixed or…
This notebook evaluates ETF features one at a time by calculating a date-by-date rank information coefficient between feature values and subsequent monthly returns. It screens financial and model-derived features over walk-forward validation dates, then…
This notebook evaluates candidate features for ranking US equities by subsequent returns. It computes daily cross-sectional information coefficients, summarizes their average with uncertainty estimates that account for serial dependence, and adjusts…
This notebook estimates the adjusted effect of a continuous ETF momentum measure on forward returns using double machine learning. It contrasts an unadjusted regression with DML estimates that control for recent and longer-term volatility, market regime, and…
This guide describes a daily US equities dataset from NASDAQ Data Link’s Wiki Prices, with adjusted OHLCV data and coverage from 1962 through March 2018. It explains how to obtain the archive with an API key, load it, filter by symbol or date, inspect its…
This notebook explains how a conditional autoencoder extends instrumented principal component analysis (IPCA): it retains the two-stage structure in which fund features map to latent factor exposures and those exposures combine with factor returns, but uses…
This analysis compares predictive models for ranking NASDAQ-100 stocks by their next 15-minute return using intraday microstructure features such as spreads, depth imbalance, signed volume, and volatility. It emphasizes selecting comparable prediction sets…
This notebook compares sklearn HistGradientBoosting, XGBoost, LightGBM, and CatBoost for predicting forward ETF returns. It measures cross-sectional information coefficient, training time, and memory use across model-complexity presets, with GPU runs…
This notebook explains how to parse IEX DEEP messages and maintain an aggregated limit order book at each price level. It extracts price-level updates, best bid and ask quotes, and trade reports, then uses the resulting data to examine spread and depth. The…