Monte Carlo Sample Size for Estimating Rare Event Probabilities
Summary
The document asks how to choose a Monte Carlo sample size for estimating a Bernoulli mean to a specified absolute precision. It presents the central limit theorem formula for sample size, using a normal quantile, the outcome variance, and the squared error tolerance. Substituting Bernoulli variance, p(1-p), produces smaller sample estimates as the event probability falls across the examples shown.
The apparent paradox comes from the precision measure: the calculation targets the same absolute error for every probability, not the same relative error. An estimate near zero can have a small absolute deviation while still being far off in percentage terms. The post contains no accepted answer or resolution, but its setup highlights why a rare-event simulation may need far more draws when the goal is reliable relative accuracy or a useful chance of observing enough successes. The displayed calculation also relies on a normal approximation, whose adequacy can be limited when the expected number of successes is small.
Key ideas
- The sample-size expression uses the variance of the simulated outcome and the square of the absolute error tolerance.
- For a Bernoulli outcome, variance depends on the success probability as p times one minus p.
- The examples target fixed absolute precision, which does not guarantee fixed relative precision across probabilities.
- Rare events can require many simulations when the goal is to observe enough successes for stable estimation.
- The central limit theorem approximation may be unreliable when the expected success count is small.
Tags
Full text
# Monte Carlo convergence sample size
# Monte Carlo convergence sample size
I'm studying Monte Carlo analysis but I find very counter-intuitive the computation of the minimum sample size in order to reach a certain level of precision. As stated in `Montecarlo methods in Financial Engineering`, such sample size comes from the CLT and it can be computed according to
$n = \frac{z^2_{\delta/2} \sigma^2}{\epsilon^2}$
My experiment is the classic bernoulli variable, with $p$ probability of success. Suppose my level of precision is $0.01$, then my $z_{0.01/2} = 2.57$.
I consider three cases.
$p=0.1\%$
$p=1\%$
$p=10\%$
With $\sigma^2 = p(1-p)$, I get
$n_{10\%} = 2.57^2 \cdot 0.1 \cdot 0.9 / 0.01^2 = 5944$
$n_{1\%} = 2.57^2 \cdot 0.01 \cdot 0.99 / 0.01^2 = 654$
$n_{0.1\%} = 2.57^2 \cdot 0.001 \cdot 0.999 / 0.01^2 = 66$
Intuitevely I would think the rarer the event, the greater the number of the simulations needed in order to approach the theoretical mean. What am I getting wrong?Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.