Assessing GPD Fit Reliability with Small Samples for Expected Shortfall
Summary
The document examines whether a small historical sample can support fitting a Generalized Pareto Distribution (GPD) for extreme-tail risk estimates. The motivating example uses monthly portfolio data to estimate Expected Shortfall at the 99% level, and asks whether a minimum sample of ten observations is adequate and how fitting error could be quantified.
One response describes repeated fits on simulated GPD data with samples of ten, then inspects the estimated shape and scale distributions. It expresses doubt that such a small sample would yield useful results, but supplies no numerical error threshold or general minimum. Another response suggests estimating the standard deviations of parameter estimates from the available data as a quality measure. It cautions that sample size alone cannot determine accuracy because estimation quality also depends on the parameter values. The discussion is illustrative and does not establish a validated sample-size rule for portfolio Expected Shortfall.
Key ideas
- The document questions whether ten observations can support reliable GPD parameter estimates.
- Repeated simulated fits can illustrate how estimates vary in small samples.
- The standard deviations of parameter estimates can serve as an uncertainty measure.
- Sample size alone does not determine estimation quality, which also depends on parameter values.
- The discussion provides no general minimum sample size or calibrated error bound.
Tags
Full text
# How many data points are required to perform a fitting of GPD?
# How many data points are required to perform a fitting of GPD?
A friend of mine told me that their firm is using Extreme Value Theory (EVT) to compute value of the Expected Shortfall 99% of a portfolio for their asset allocation process. To do so, they try to fit the parameters of the Generalized Pareto Distribution using the historical sample of their portfolio, and they use a formula to get the value of $ES_{0.99}$.
However, I they are using monthly data points, which means that the size of the sample is pretty (very) limited. When I asked what the minimum number of point required to performed the fitting was, I was told that their algorithm was asking for a minimum of 10 points.
I am wondering
- if that's enough?
- how could we compute some kind of quantitative value indicating "how wrong" the fitting might be given the give size of the sample $N$? (this method could be specific to GPD, but of course a generic approach would be great).
## Answer by Owe Jessen (score 2)
https://quant.stackexchange.com/a/3365
Playing around with some random numbers from the GPD, I'm not convinced that a sample of 10 will give anything like useful results.
The code for generation was
```
require("gPdtest")
reps = 10000
ValuesAct = rgp(reps, 0.5,0.5)
plot(density(ValuesAct))
quantile(ValuesAct, probs=0.999)
loops = 10000
results = matrix(rep(NA,2*loops), ncol = 2)
colnames(results) = c("shape", "scale")
for(i in 1:loops)
{
ValuesSample = sample(ValuesAct,10,replace=F)
results[i,] = gpd.fit(ValuesSample, "amle")
}
plot(density(results[,2]))
lines(density(results[,1]), col = "green")
abline(v=0.5)
legend("topright", c("Shape", "Scale"), col = c("green", "black"), lwd=1)
```
## Answer by Bob Jansen (score 0)
https://quant.stackexchange.com/a/3137
Given a sample you can estimate the parameters and calculate the standard deviation of these estimations. You could use this as a measure of quality.
Creating a measure given only the sample size $N$ seems very hard to me since the quality of the estimation may depend on the values of the parameters.Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.