This module presents post-training checks for time-series generators, based on the TimeGAN evaluation approach. It measures utility by training a recurrent predictor on synthetic sequences and testing it on real data. One mode follows the paper’s setup by…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
715 documents
This notebook uses Optuna to tune XGBoost, LightGBM, and CatBoost on a time-split firm-characteristics dataset. Each library’s search treats the loss function, either mean squared error or mean absolute error, as a categorical hyperparameter alongside model…
This benchmark compares pandas and Polars on operations found in financial data pipelines, including rolling features, group calculations, window transformations, filtering, joins, lazy scans, memory use, and string handling. It generates shared synthetic…
This notebook applies TabM, a tabular neural-network approach, to foreign-exchange pair prediction rows without treating the data as a sequence. It uses shared runner infrastructure to fit preprocessing within each training fold, save declared weight…
This document explains how to turn weekly Commitment of Traders reports into futures positioning features. It outlines the report categories for financial futures and physical commodities, describes how net positions reflect different participant roles, and…
This notebook describes a read-only comparison of validation predictions from several model families on a US equities panel. It first checks that each predefined prediction set is complete, then assesses cross-sectional ranking with the information…
This document explains how hidden Markov models infer unobserved market regimes from returns and recent volatility. It first sets two transparent benchmarks: a volatility index threshold for stress and price relative to a long moving average for trend. It…
This notebook implements a univariate forecast of SPY daily closing prices using raw PyTorch, sktime, and Darts. It compares the practical experience of each interface, including implementation effort, installation constraints, and combined fit-and-predict…
This notebook explains a family of ETF models that represents returns through shared latent directions and estimates how fund features map to exposures. It distinguishes five approaches: unconditional principal components, instrumented PCA with a linear…
This analysis compares predictive, latent-factor, and causal models for a cross-sectional S&P 500 stock strategy using weekly forward returns. Its feature set combines equity momentum and volatility measures with option information such as implied-volatility…
This notebook describes reconstructing a single-symbol, single-day NASDAQ limit order book from message-by-order ITCH data. It processes add, execute, cancel, delete, and replace messages while maintaining each live order reference and its remaining shares.…
This notebook profiles an annual-report corpus before indexing it for financial research. It highlights four data issues that row counts and null checks do not reveal: the gap between fiscal year-end and public filing, duplicate records when one filing…
This notebook assesses whether an LSTM can use the ordering of ETF feature histories to improve on flat-feature linear and gradient-boosting models. It resolves the declared sequence population against current data, checks eligible funds and fund-date rows,…
This notebook uses a synthetic asset panel to demonstrate Instrumented PCA, where factor loadings depend linearly on characteristics observed before returns. Alternating least squares estimates the characteristic-to-loading map and realized factors. Because…
This notebook explores ARIMA as a feature generator for models that may use its output alongside other predictors. It explains how ACF and PACF plots can suggest model orders, how information criteria can compare candidate orders on training data, and why…
This reference describes academic factor-return datasets from the Fama-French library and AQR for use in strategy analysis and factor modeling. Fama-French offerings include market, size, value, profitability, investment, and momentum series, alongside…
This research-agent record considers whether the Federal Reserve will raise the upper bound of its target rate during 2026. It contains a market price, search traces, and agent probability estimates. The first rationale favors a hike, citing inflation risks…
This notebook explains how double machine learning (DML) estimates whether FX momentum affects future returns after adjusting for configured confounders. It distinguishes this intervention question from prediction: predictive models are compared by…
This notebook presents lightweight falsification diagnostics for feature triage, explicitly distinguishing mechanism consistency from causal identification. It first scans ETF features across forward-return horizons with multiple-testing correction, then…
This notebook aggregates registry results for OLS, Ridge, Lasso, and ElasticNet across nine case studies. It selects complete validation results at each study's primary label, compares mean daily cross-sectional rank information coefficients and HAC…
This notebook trains Skip-gram Word2Vec on labeled financial news sentences and explains how its vectors represent words that occur in similar contexts. It describes the roles of vector size, context window, minimum frequency, prediction mode, negative…
This notebook evaluates ETF features one at a time by calculating a date-by-date rank information coefficient between feature values and subsequent monthly returns. It screens financial and model-derived features over walk-forward validation dates, then…
This notebook evaluates candidate features for ranking US equities by subsequent returns. It computes daily cross-sectional information coefficients, summarizes their average with uncertainty estimates that account for serial dependence, and adjusts…
This notebook shows how to align macroeconomic observations with the dates traders could actually have known them. It distinguishes the period a value measures from its publication date, estimates release dates from period length and agency lag schedules,…