Skip to content
All library documents

Choosing Loan-Level or Aggregated Data for MBS Prepayment Models

Article Quant Q&A · Author: Jojo

Summary

The document considers how to model monthly prepayments for a mortgage-backed securities sector when the dataset contains millions of loans over several years. It asks whether to aggregate loan observations by month using weighted averages of explanatory variables, or retain a longitudinal or panel structure. The central issue is a trade-off: aggregation can lose information and introduce bias, while the appropriate data arrangement depends on the application, available features, and computational resources.

The answer gives no specific correction to apply to predictions or regression parameters after aggregation, and it presents no empirical comparison. It advises that, for a simple linear regression, the choice of aggregation level may matter less for predictive or explanatory performance than specifying the model well and selecting useful explanatory variables. That is a general guideline, not a guarantee; the document does not define weighting, model diagnostics, or how to handle loan-level dependence over time.

Key ideas

  • Aggregating loan observations by month can discard information and create aggregation bias.
  • The choice between pooled, longitudinal, and panel data depends on the modeling task and available resources.
  • The answer frames aggregation as a trade-off that has no universally correct resolution.
  • For a simple linear regression, feature selection and model specification may matter more than aggregation level.

Tags

Full text
# How to set up data for understanding drivers of prepayments


# How to set up data for understanding drivers of prepayments












I would like to understand the drivers of prepayment of a certain sector of MBS. I have some explanatory variables that I think would explain the actual CPR's and want to model the prepayments through a simple linear regression. I have millions of loans and several years worth of monthly data. To my understanding, I need to pool this data together for each timestamp (month) before running this regression. What I wanted to understand is, when grouping the data by time and taking the weighted averages across the explanatory variables, I end up to some extent loosing information, so is there other ways the data for prepayments is put together aside from grouping in this manner? Is it fine to do just do this grouping and then running the regression, and are there any adjustments made to predictions/parameters after the regression is run to account for the grouping? I guess I'm just wondering if the data is usually set up as longitudinal (which I am trying to do) or panel data?

## Answer by Sharad (score 2, accepted)

https://quant.stackexchange.com/a/58591

This is a complex decision which involves the trade-off between aggregation bias and measurement error. For example, see link. In general, there is no cut-and-dried answer -- the appropriate level to group at depends upon the specific application for the model, the feature set, and access to computational resources, among other factors.

Given that you are attempting to model prepayments using a simple linear regression, your choice of the level to group at is likely to have less of an impact on the predictive/explanatory power of your model than your choice of model specification and the explanatory variables you choose to include.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.