This notebook develops a financial feature matrix for a cross-asset ETF momentum hypothesis: assets with stronger relative performance may continue to outperform over the following month. It combines trailing returns at several horizons, risk-adjusted…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
742 documents
This document describes refitting the configuration selected by earlier validation stages on all eligible pre-2021 history, then generating predictions for the 2021 holdout. It derives the training interval from the declared evaluation window, label buffer,…
This notebook tests whether gradient-boosted trees improve cross-sectional ranking across currency pairs beyond a penalized linear model. The FX universe contains pairs sharing currencies, so observations are dependent: a move in one currency affects…
This utility module supports deep learning workflows for financial time series across multiple assets. It resolves dataset aliases and loads canonical case study data, then creates sliding-window sequences independently for each symbol. The sequence…
This notebook compares two locally run, open-weight embedding models on passages from recent 10-K filings and a fixed set of financial research queries. It builds two-sentence passages, embeds the same documents and queries with each model, and evaluates…
This notebook constructs price-derived features for a broad US equities panel, including momentum, moving averages, and volatility measures. It is designed to rank stocks against one another, using a tradability screen, per-symbol rolling calculations, and…
This notebook compares three ways to orchestrate a four-phase forecasting workflow: direct composition in Python, a role-prompted CrewAI version, and a LangGraph version that delegates to the same specialist classes as the native implementation. It examines…
This notebook builds minute-level features from NASDAQ-100 quote and trade data to study short-horizon price pressure. It treats normalized order-flow imbalance as the main signal candidate and uses spread, depth, price impact, off-exchange trading,…
This document describes a ledger for applying funding cash flows to perpetual futures positions during a backtest. At each funding timestamp, it uses the position’s signed quantity, the current mark, any contract multiplier, and the funding rate to calculate…
This notebook checks whether four-hour spot FX data can support a daily cross-sectional strategy that ranks currency pairs using momentum and carry. It tests whether the declared instruments have prices at each decision point, whether the universe represents…
This notebook evaluates stop-loss, trailing-stop, and fixed-duration exits as overlays on CME futures strategies. For each prediction horizon, it applies configured rules to the strongest validation-Sharpe parent selected from prior signal and allocation…
This notebook screens financial and model-based features for their ability to rank stocks by a forward return. It computes daily cross-sectional information coefficients, estimates uncertainty while accounting for serial dependence, adjusts significance for…
This notebook compares three linear forecasters, a Transformer encoder, and two parameter-free forecasts on daily SPY returns. Linear models map a historical window to a multi-day forecast; variants first separate a smoothed component or account for the last…
This notebook configures an NLinear forecasting run for FX pairs. The model uses a fixed consecutive lookback and subtracts the last observed level, directing its fit toward changes over that window. It resolves the lookback, normalization, device, folds,…
This guide explains how to turn hourly continuous futures data into daily bars aligned to CME trading sessions. Because a session ends at 4 PM Central Time, bars from Sunday evening belong to Monday's session, and bars after the close generally count toward…
This code defines safeguards for reproducible cross-validation, eligibility tracking, and fold-scoped temporal features. It normalizes fold boundaries, compares requested folds with the boundaries used to create temporal artifacts, and rejects incompatible…
This notebook recasts prediction of 21-session ETF forward returns as a binary task: positive returns are labeled up, and all others down. It fits L2- and L1-regularized logistic regression using chronological walk-forward folds with a purge gap. Scaling is…
This notebook tests whether an LSTM can extract temporal structure from ETF feature histories that flat-feature models may miss. It resolves the declared sequence population against available features, labels, entities, and walk-forward folds before fitting.…
This notebook defines close-to-close forward returns over two trading horizons for a cross-section of ETFs, with each horizon measured from an adjusted close to the close a fixed number of sessions later. It explains why labels are built on complete symbol…
This tutorial derives the Kelly fraction for a binary wager by maximizing expected logarithmic wealth growth, then extends the idea to continuous returns and a multi-asset portfolio. It uses symbolic differentiation and simulations with shared coin-toss…
This notebook evaluates four signals derived from news text: weighted surprise, average sentiment, sentiment momentum, and article coverage. It uses forward returns prepared by an earlier feature-building step, then calculates a daily cross-sectional…
This chapter treats transaction costs as a constraint throughout strategy research and deployment, from factor evaluation and backtesting to portfolio construction, risk oversight, and production monitoring. It distinguishes explicit fees, implicit spread…
The notebook demonstrates tuning LightGBM for ETF return prediction with Optuna’s TPE sampler. It searches tree structure, sampling, and regularization settings, using early stopping to choose the boosting rounds and a custom callback to report…
This chapter presents a research workflow for using predictive models in trading, where stable out-of-sample forecasts may matter more than unbiased coefficient estimates. It covers regularized regression methods such as Ridge, LASSO, and Elastic Net, along…