Modeling Trend and Seasonality in Loan Default Rates
Summary
The document describes a backtesting problem in which a model for loan non-performing-loan rates assumes observations have a stable mean, but the rates move through trends or cycles. In that setting, a large sample does not necessarily bring the observed mean close to the assumed long-run mean, so a confidence interval based on independent, identically distributed observations can make the model appear to fail. The cycles may vary in duration, making them difficult to predict directly.
One answer proposes decomposing the series into cycle, seasonality, and residual components, modeling the more identifiable seasonal pattern first, then fitting a distribution to the residuals. Another suggests estimating and subtracting a mean-reverting trend with an Ornstein–Uhlenbeck-style process, using regression to estimate its parameters. These are candidate approaches rather than a worked comparison: the document provides no empirical results, and both methods depend on whether the chosen trend or seasonal structure captures the observed series without using future information.
Key ideas
- A stable-mean backtest can fail when observations are trending or cycling over the sample period.
- A time-series model can separate cyclic, seasonal, and residual components before modeling residual variation.
- Seasonality may be easier to estimate when it follows a recurring calendar pattern.
- A mean-reverting process can be used to estimate and subtract a trend component.
- Trend and seasonal adjustments require validation against the series’ actual structure.
Tags
Full text
# how to make a distribution model tolerable of trend?
# how to make a distribution model tolerable of trend?
I'm building an model on different loans' NPL rate. The problem is NPL rates are always affected by the market. When NPL rates move in trend, my model will fail the back-testing.
Assuming $x(t)$ is a random variable that distributed among $[-1, 1]$, with the mean $\mu = 0$ and a standard deviation $\sigma$.
When the sample size $n$ is big, the distribution of observed mean $\bar x$ will be ~ $N(0, \sigma^2/n)$. The back-testing confidence interval for $\bar x$ is $[- z_{\frac{\alpha}{2}} \frac{\sigma}{\sqrt{n}} , z_{\frac{\alpha}{2}} \frac{\sigma}{\sqrt{n}} ]$. The model always passes the backtest.
Now, troubles come when $x(t)$ has some trend. Let's say $x(t)$ becomes $x(t) = \sin (t + random(t) )$. Here $x(t)$ still follows the distribution, but when $x(t)$ moves near $+1$, the samples' mean will be around $+1$, increasing the sample size will not bring the $\bar x$ near to $0$, the model fails the backtesting.
Now my problem is, NPL's trend is hard to predict, the cycle sometimes are 6 months, sometimes are 2 years. My NPL model always fails the back-testing because of trend.
Any suggestion please?
## Answer by Richi Wa (score 1)
https://quant.stackexchange.com/a/8856
In my experience with forecasting, you could try a model of the form $$ X_ t = cycle_t + seasonality_t + residuum_t. $$ Sometimes it is hard to find the cycle but the seasonality could be doable if it has some natural structure (something happening in a certain month each year e.g.). Rob Hyndman explains all these things (and provides an R package) in his free online book. You could have a look at chapter 6. After having calculated the first 2 components you could model the distribution of the residuum.
## Answer by user7056 (score 0)
https://quant.stackexchange.com/a/8603
you could calculate and subtract the trend.
dx=h(m-x)*dt+s*dz x_(t) - x_(t - 1) = m (1 - exp(- h Dt)) + (exp(- h Dt) - 1) x_(t - 1) + e_t error e_t normally distributed with
(s_e)^2 = [1 - exp(- 2 h)] (s^2)/2h
Because when Dt->0
```
x_t - x_(t - 1) ~h(m-x_(t - 1))dt + et
```
you calculate the parameters h and m from the mean regression after that....?Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.