Partial Sample Regression Using Historical Relevance
Summary
This document explains history-weighted, or partial sample, regression as a way to make predictions from observations judged relevant to a new input. It defines similarity using negative Mahalanobis distance and informativeness by how far an observation lies from the training data mean. Their combination ranks historical observations; a selected top-quantile subsample is then used in a relevance-weighted prediction.
The article connects the method to ordinary least squares: when the full sample is retained, the prediction reduces to the OLS result. It also distinguishes the approach from weighted least squares and from fitting a separate regression on the selected subset, since relevance is calculated using the full-sample covariance. An example forecasts quarterly US GDP from economic indicators and notes that full-sample regression matches OLS. The method depends on a suitable covariance estimate and relevance selection; the article gives no broad performance comparison or evidence that selecting a subsample improves forecasts. Feature rescaling can help numerical conditioning without changing the prediction.
Key ideas
- Mahalanobis distance measures how similar two observations are relative to feature covariance.
- Informativeness increases as an observation lies farther from the training data mean.
- Partial sample regression ranks observations by relevance and predicts from a selected subset.
- Using the entire training sample makes the prediction equivalent to ordinary least squares.
- The full-sample covariance is retained even when only a subset of observations is selected.
Tags
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.