This notebook explains a principal-component baseline for forecasting CME futures returns. PCA is fitted to the training fold’s product-return panel, so its components represent co-movement directions across contracts rather than compressed carry, momentum,…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
87 documents
This analysis explains how CME futures data is organized into products, expiring contracts, and volume-rolled continuous series. It uses E-mini S&P 500 data to show how individual contracts have finite trading windows that overlap around a roll, while a…
This notebook compares equal-weight, positive-score-weighted, and conformal-width-weighted allocations for top-ranked ETFs and CME futures. It estimates entity-specific split-conformal interval widths from residuals in earlier validation folds, excluding…
This notebook uses registered walk-forward GBM predictions for ETFs and CME futures to compare equal weighting, inverse conformal-width weighting, and weighting proportional to positive prediction scores. It calibrates entity-specific Mondrian…
This tutorial constructs a continuous futures series from individual ES contracts. It identifies the front month using trading volume, with a monotonic roll constraint, and compares volume-based switching with calendar-based timing. It explains why outright…
This notebook isolates results for the LEAN engine from an audit comparing retained real-strategy runs with matching ML4T Backtest profiles. It identifies four supported workloads: ETF allocation, crypto perpetual funding, USD-quoted FX allocation, and a US…
This notebook uses real ETF data to assess whether a momentum signal remains useful under reasonable changes to its lookback, market regime, and implementation. It computes cross-sectional information coefficients between momentum and forward returns, then…
This notebook compares ways to measure volatility and develops heterogeneous autoregressive (HAR) volatility models alongside roughness analysis. It uses intraday returns to estimate session realized variance, distinguishes intraday movement from overnight…
The notebook builds prediction intervals around gradient boosting forecasts with split conformal prediction. It fits a model on an earlier training segment, measures absolute residuals on a later calibration segment, and uses a finite-sample order statistic…
This audit compares trading engines replaying the same frozen model targets and content-addressed historical inputs across ETF allocation, futures, crypto perpetual funding, foreign exchange, and US equities. It evaluates execution parity through matching…
This reference summarizes listed contract months for 35 CME futures products across equity indexes, Treasuries, energy, metals, currencies, interest rates, agriculture, livestock, and crypto. It explains the exchange’s month-code system and distinguishes…
This notebook uses double machine learning to estimate whether futures carry predicts subsequent returns independently of volatility, momentum, and cross-sectional carry rank. It residualizes both carry and returns against these confounders with flexible…
This case study lays out a research pipeline for crypto perpetual futures, treating funding payments exchanged between long and short positions at regular settlements as a potential return source. It describes data and model stages from label construction…
The notebook explains how to read Kalshi’s binary Federal Reserve rate contracts as probabilities and how to turn a ladder of rate thresholds into an implied distribution. It stresses that the feed contains YES bids, so prices are lower bounds affected by…
This document compares alternative position-sizing methods for CME futures strategies selected from a baseline ranked by equal-weight validation Sharpe. Equal weighting treats every selected product alike, although futures contracts can have very different…
This notebook uses double machine learning to estimate whether futures carry has an effect on subsequent returns after adjusting for volatility, momentum, and cross-sectional carry rank. It distinguishes causal explanation from predictive performance: a…
This notebook applies a neural stochastic discount-factor model to CME futures. In asset-pricing terms, a stochastic discount factor is a random variable whose product with each asset's return has the same expected value under no arbitrage. The model…
This analysis checks whether settlement data can support a weekly, cross-sectional CME futures strategy before any model is fit. It tests whether enough products are quoted on rebalance dates to fill the intended long and short portfolios, estimates spread…
This notebook uses Optuna to tune LightGBM return models while considering both cross-sectional information coefficient (IC) and prediction turnover. A single-objective search maximizes IC for comparison; an NSGA-II multi-objective search identifies…
This analysis explains how to read cross-sectional information coefficients (ICs) for CME futures models and how they relate to later portfolio tests. It computes each date’s rank correlation between predicted and realized returns, then averages those…
This notebook checks whether the data can support a weekly, cross-sectional futures strategy before any model is fitted. It describes a design that ranks CME products, takes long positions in the highest-ranked contracts and short positions in the lowest,…
This notebook screens candidate features for a CME futures case study against forward returns. For each feature, it computes cross-sectional rank correlations between feature values and subsequent returns, summarizes the correlation over dates, adjusts…
This notebook studies gradient boosting for cross-sectional prediction of CME futures returns. It varies tree capacity through leaf-count profiles and compares squared-error, absolute-error, and Huber objectives, which differ in how strongly extreme…
This module defines a reader-facing workflow for CME futures research, covering model requests, prediction and backtest execution, candidate pools, and final selection. It reads the configured label sweep from a shared setup declaration, preserves its order,…