This notebook explains PCA as a baseline latent-factor model for forecasting ETF returns. It decomposes a training panel of forward returns into leading directions and estimates each fund’s exposure and factor premia. The estimator ignores the available…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
715 documents
This notebook demonstrates fine-tuning three transformer checkpoints for three-class financial sentence sentiment and evaluating them with accuracy, macro F1, and confusion matrices. It uses a stratified train, validation, and test split so class imbalance…
This notebook estimates the adjusted effect of a continuous ETF momentum measure on forward returns using double machine learning. It contrasts an unadjusted regression with DML estimates that control for recent and longer-term volatility, market regime, and…
This notebook explains how a conditional autoencoder extends instrumented principal component analysis (IPCA): it retains the two-stage structure in which fund features map to latent factor exposures and those exposures combine with factor returns, but uses…
This analysis compares predictive models for ranking NASDAQ-100 stocks by their next 15-minute return using intraday microstructure features such as spreads, depth imbalance, signed volume, and volatility. It emphasizes selecting comparable prediction sets…
This notebook compares sklearn HistGradientBoosting, XGBoost, LightGBM, and CatBoost for predicting forward ETF returns. It measures cross-sectional information coefficient, training time, and memory use across model-complexity presets, with GPU runs…
This notebook maps a family of ETF return models that infer common latent directions in a panel, with features used to estimate fund exposures. It distinguishes five approaches: unconditional principal components; instrumented PCA with a linear…
This notebook shows why searching across many signals or strategies makes the top observed result look stronger than its underlying predictive value. A simulation uses factors with no true information to illustrate how selecting the largest information…
This notebook explains how to turn a released, cross-sectionally ranked US firm characteristic panel into a keyed feature matrix for a factor study. It groups inputs into declared families, combines characteristics and interactions without fitting…
This US equities feature study explains how to generate features from estimated models without allowing future data into earlier observations. Its estimation schedule uses a history burn-in, fits parameters only on data before each output block, then…
This notebook applies TSMixer to ETF sequences using one globally shared function across funds. Each example contains an individual fund's history and covariates; temporal and feature interactions are learned across the panel, but one fund's observations are…
This notebook describes a stochastic discount factor (SDF) model for pricing the cross-section of US firm returns. Instead of first estimating factors and then applying them, it directly learns firm-month weights under a no-arbitrage condition: discounted…
This notebook analyzes reconstructed NASDAQ limit order books to describe intraday spreads and top-of-book depth, then examine whether order-flow imbalance is associated with subsequent bucket returns. It expresses spreads in basis points to compare stocks…
This notebook queries backtest registries across nine case studies and assembles comparable tables for downstream analysis. It organizes results by asset class and data frequency, then records performance across signal, allocation, cost, and risk stages,…
This notebook evaluates four news-derived signals—weighted surprise, average sentiment, sentiment change, and article coverage—against forward stock returns. It computes a daily cross-sectional Spearman information coefficient, summarizes its mean,…
This document describes fitting temporal convolutional networks to NASDAQ 100 minute level microstructure features. Causal convolutions prevent a prediction from using later observations, while dilation lets successive layers capture patterns over multiple…
This notebook specifies and executes a double machine learning analysis of the effect of the variance risk premium on short-option returns through expiry. Before execution, it resolves the treatment, outcome, confounders, timing, nuisance model, temporal…
This analysis compares predictions from several model families trained on monthly US stock characteristics to forecast next-month returns. It focuses on cross-sectional information coefficient, which measures how well a model ranks stocks within each month.…
This case study fits regularized linear models to returns from short at-the-money straddles held to expiry. The trade collects call and put premiums, giving it a capped maximum gain but potentially very large losses when the underlying moves sharply.…
The document explains how an experiment registry can track a model from its training configuration through predictions to backtest results. Each stage receives an identifier derived from a canonicalized specification, allowing repeated identical runs to…
This notebook compares methods for discovering relationships among a panel of ETF returns: NOTEARS for contemporaneous linear directed acyclic graphs, VAR-LiNGAM for lagged and instantaneous structure, PCMCI for conditional-independence links, and Granger…
This notebook demonstrates tuning LightGBM hyperparameters with Optuna's TPE sampler, using cross-sectional information coefficient as the objective. It combines early stopping to choose the number of boosting rounds with a custom pruning callback that…
This notebook studies how ridge, lasso, and elastic net behave when a crypto perpetuals feature matrix measures one economic quantity—the premium—many different ways. Premium levels, changes, volatility, standardized positions, ranks, and related funding…
This notebook brings together five latent-factor approaches for modeling the S&P 500 options case study’s equity return cross-section. PCA estimates common movements from returns alone; IPCA maps characteristics to exposures linearly; a conditional…