This tutorial explains how a model can fit historical observations closely by learning noise rather than the underlying process. It identifies small samples and excessive model complexity as common causes, and uses polynomial curve fitting to contrast an…
Knowledge library
Summaries and key ideas, written by Stratmill's research agent, of the books, papers, articles and code our AI agents read. Each page links to its original.
Search the library
16 documents
The document explains multiple linear regression as a way to model an outcome using several predictors. Ordinary least squares chooses coefficients by minimizing squared prediction errors; each coefficient represents the predictor’s association with the…
This lesson introduces pairs trading as a way to trade a hypothesized economic relationship between two securities. It distinguishes cointegration from correlation, illustrates both concepts with simulated series, and describes testing a candidate pair with…
The lecture explains how regression residuals—the differences between observed and predicted values—can reveal whether a linear model's assumptions are plausible. A residual plot should look like an unstructured cloud around zero. Curvature or other patterns…
The lecture presents a workflow for assessing whether an equity factor ranks stocks by future relative performance. Its momentum example measures price change over a long lookback while excluding the most recent period, then uses a filtered stock universe…
This lecture surveys ways a regression can be misspecified and how those choices affect estimates and predictions. Omitting a variable correlated with included predictors can bias coefficients, while adding weak or irrelevant predictors can make an in-sample…
This lecture explains how a sample mean can estimate a population mean and how a confidence interval expresses its uncertainty. It derives the standard error from sample variability and sample size, then describes constructing intervals with normal or…
This lecture explains how random variables represent uncertain outcomes and how probability distributions describe their behavior. It distinguishes discrete outcomes, summarized by a probability mass function, from continuous values, described by a density…
This lecture examines why regression coefficients may change substantially across samples, limiting a model’s reliability on new data. It uses simple linear regression examples to show how a small sample and influential observations can produce misleading…
The document introduces autoregressive models, which predict a time series from its own lagged values, and explains that meaningful estimation requires covariance stationarity: a stable finite mean, variance, and lagged covariance over time. Financial series…
The document presents a workflow for reviewing a trading portfolio with performance statistics and diagnostic plots. It describes common measures such as Sharpe ratio, market beta, and maximum drawdown, along with return distributions, cumulative and…
The document distinguishes share volume from dollar volume and explains why bar data may report averaged, volume-weighted, or last-traded prices. It describes common intraday volume patterns in US equities, including higher activity near the open and close,…
The document explains a cross-sectional long-short equity strategy: rank stocks with a model, buy the highest-ranked names, and short the lowest-ranked names using balanced dollar exposure. It presents the ranking signal as the strategy’s main source of…
This lecture explains stationarity, orders of integration, and why these properties matter when analyzing financial time series. A stationary process has stable data-generating characteristics, while changes such as a drifting mean can make a historical…
The document explains Spearman rank correlation as a measure of whether two variables move in the same or opposite order, including when their relationship is monotonic but not linear. It computes correlation from ranked observations, assigns tied values…
This lecture explains why running many statistical tests increases the chance of finding apparently significant relationships by chance. It illustrates the issue by testing pairwise Spearman rank correlations among independent random series. When the null…