The document explains how a stochastic discount factor (SDF) estimates a pricing kernel that should price every asset, rather than estimating common return factors. It describes adversarial training: one network proposes the discount factor while another…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
129 documents
This notebook uses Optuna to tune XGBoost, LightGBM, and CatBoost on a time-split firm-characteristics dataset. Each library’s search treats the loss function, either mean squared error or mean absolute error, as a categorical hyperparameter alongside model…
This notebook explains a family of ETF models that represents returns through shared latent directions and estimates how fund features map to exposures. It distinguishes five approaches: unconditional principal components, instrumented PCA with a linear…
This notebook uses a synthetic asset panel to demonstrate Instrumented PCA, where factor loadings depend linearly on characteristics observed before returns. Alternating least squares estimates the characteristic-to-loading map and realized factors. Because…
This reference describes academic factor-return datasets from the Fama-French library and AQR for use in strategy analysis and factor modeling. Fama-French offerings include market, size, value, profitability, investment, and momentum series, alongside…
This notebook explains PCA as a baseline latent-factor model for forecasting ETF returns. It decomposes a training panel of forward returns into leading directions and estimates each fund’s exposure and factor premia. The estimator ignores the available…
This notebook explains how a conditional autoencoder extends instrumented principal component analysis (IPCA): it retains the two-stage structure in which fund features map to latent factor exposures and those exposures combine with factor returns, but uses…
This notebook maps a family of ETF return models that infer common latent directions in a panel, with features used to estimate fund exposures. It distinguishes five approaches: unconditional principal components; instrumented PCA with a linear…
This notebook shows why searching across many signals or strategies makes the top observed result look stronger than its underlying predictive value. A simulation uses factors with no true information to illustrate how selecting the largest information…
This notebook explains how to turn a released, cross-sectionally ranked US firm characteristic panel into a keyed feature matrix for a factor study. It groups inputs into declared families, combines characteristics and interactions without fitting…
The document describes how a trading research pipeline assesses whether latent-factor model fits completed in a usable state. For models trained by gradient descent, it checks that the final recorded training objective is finite; for the stochastic discount…
This notebook describes a stochastic discount factor (SDF) model for pricing the cross-section of US firm returns. Instead of first estimating factors and then applying them, it directly learns firm-month weights under a no-arbitrage condition: discounted…
This notebook evaluates four news-derived signals—weighted surprise, average sentiment, sentiment change, and article coverage—against forward stock returns. It computes a daily cross-sectional Spearman information coefficient, summarizes its mean,…
This analysis compares predictions from several model families trained on monthly US stock characteristics to forecast next-month returns. It focuses on cross-sectional information coefficient, which measures how well a model ranks stocks within each month.…
This notebook brings together five latent-factor approaches for modeling the S&P 500 options case study’s equity return cross-section. PCA estimates common movements from returns alone; IPCA maps characteristics to exposures linearly; a conditional…
This notebook screens financial and model-based features for their ability to rank stocks by a forward return. It computes daily cross-sectional information coefficients, estimates uncertainty while accounting for serial dependence, adjusts significance for…
This notebook evaluates four signals derived from news text: weighted surprise, average sentiment, sentiment momentum, and article coverage. It uses forward returns prepared by an earlier feature-building step, then calculates a daily cross-sectional…
This notebook explains how to decompose ETF returns and risk using CAPM and Fama–French factor regressions. It estimates full-sample exposures with heteroskedasticity and autocorrelation robust standard errors, tracks changing betas with rolling windows, and…
This document explains a supervised autoencoder factor model for predicting stock returns in an equity option analytics research setting. Unlike PCA, IPCA, and an unsupervised conditional autoencoder, its training objective combines reconstruction of the…
This notebook studies linear prediction models for a cross-section of CME futures products using feature columns grouped into related families, including carry, momentum, volatility, and rolling risk measures. Because columns within a family are often…
This notebook presents a conditional autoencoder for equity returns in which a neural network maps stock characteristics to nonlinear factor loadings, while another network extracts contemporaneous latent factors from characteristic-managed portfolio…
This notebook examines whether published investment factors offer credible return premia and diversify one another. It draws on long-history and cross-asset series from AQR alongside Fama-French equity factors, using serial-correlation-aware statistics and a…
This notebook implements a daily educational adaptation of an adversarial stochastic discount factor model. A portfolio-weight network constructs a factor from next-day excess returns using characteristics and market state available at the prior close. An…
This notebook runs a fixed equity-characteristics strategy on a reserved holdout period using predictions and a portfolio allocator selected earlier. It keeps the configuration, position sizing, concentration, rebalance cadence, and cost assumption…