Skip to content
All library documents

Using Empirical Covariance and Random Draws in Monte Carlo Studies

Article Quant Q&A · Author: JoseMM

Summary

This answer introduces Monte Carlo simulation as a way to study the sampling behavior of an estimator, especially when its distribution is not available analytically. It illustrates the idea by repeatedly drawing samples from a known normal distribution and calculating statistics such as the mean and median. Repeating the experiment produces an empirical distribution of those statistics, which can be used to assess estimation error under specified assumptions.

For the fund study in the question, the response interprets the empirical covariance matrix as one estimated from observed sample data, and describes drawing from a multivariate normal distribution using those estimated parameters. It mentions a covariance function as a common software tool. The example is pedagogical and does not fully reconstruct the cited paper’s simulation design; it also leaves the rationale for a later choice about fund means uncertain. Distributional assumptions and implementation details therefore need to be checked against the original study.

Key ideas

  • Monte Carlo simulation estimates a statistic’s sampling distribution through repeated simulated samples.
  • An empirical covariance matrix is estimated from observed sample data.
  • A multivariate normal model can use estimated means and covariance as simulation inputs.
  • Simulation conclusions depend on the chosen assumptions and design.

Tags

Full text
# Help Setting a Monte Carlo Simulation


# Help Setting a Monte Carlo Simulation












I am trying to replicate the steps of the Barras, Scaillet, Wermer(2010) paper for a Monte-Carlo Simulation. More specifically the steps in Appendix B.1 (Attached image). I have so far done the regressions for all the funds, sampled 1,400 funds, and adjusted the alphas as the authors did. But I am having trouble understanding 2 things (I am really new to this, sorry)

- I just don't understand what they mean by: 'we proxy $$\Sigma_{F}$$ by its empirical counterpart' (highlighted in yellow) How do I get that covariance matrix?

- And when they say, in the last paragraph, that they randomly draw $$\epsilon_t$$ 384 times, is it just whatever value comes from putting a Normal(0,0.021) since they are giving the sigma equally for all the cases?

I want to understand this before moving on to the Monte-Carlo simulation, which I have never done before. Also, if this is not the appropriate forum or you recommend something else I highly appreciate any help.

## Answer by Dave Harris (score 1, accepted)

https://quant.stackexchange.com/a/60888

There are three reasons to perform Monte Carlo simulations in statistics. The first, as used in this paper, is to test the performance of estimators when an analytic solution does not exist. The second is to construct scenarios for the future to determine how well fit estimators are. The third is in Bayesian statistics to determine the value of the denominator in Bayes Rule by substituting Monte Carlo methods for analytic integration. The Bayesian version requires additional steps because the draws need to be Markovian.

Imagine you knew with certainty that you were drawing from a Gaussian distribution centered on zero with unit variance. You could draw samples, of size N, to estimate the distribution of some statistic such as the sample mean, variance, median, 23rd quartile and so forth.

If you were to perform enough samplings, you would end up with estimates of the probability that some estimator will be some distance away from its true value.

The pseudo-code might look like this:

```
initialize x[30] Real;
initialize y[1000] Real;
for i=1:1000;
     j=1:30;
           x[j]=random_normal(0,1);
     next j;
     y[i]=average(x[1:30])
 next i;
```

The output would converge to the z score as $i$ became large enough.

That is a Monte Carlo simulation.

In R code, although the graphics are inelegant, it is:

```
rows<-30
columns<-100000
x<-matrix(rnorm(rows*columns),nrow = rows)
y<-apply(x,2,mean)
z<-apply(x,2,median)
plot(density(y),main = "Solid Line Average, Dotted Line Median Sampling Distribution 
Estimate")
lines(density(z),type = "p")
print(summary(as.vector(x)))
print(summary(y))
print(summary(z))
```

Sorry, it is inelegant, but it lets you see what they are doing but in the one dimensional case. The above shows the distribution of sample means, where the sample size is 30, versus the distribution of the sample medians.

Of course, both of these have analytic solutions. That is the z-table, in essence, except that $$\sigma_{median}=1.253\frac{\sigma}{\sqrt{n}},$$ because it is less efficient an estimator.

So what they are doing in their simulation is treating $\Sigma_F$ as precisely equal to the sample statistics. They are then randomly drawing hundreds of samples from a multi-dimensional normal distribution. If they did that with sample means equal to zero, then that would be the empirical estimate of the distribution under the null hypothesis.

You get the empirical covariance matrix by estimating the components using the sample statistics. In R it is estimated with the cov() function, unless you are doing a regression, then it comes as an implementation of the function. Its language varies with the tool.

As to why they draw uniformly on the means later in the treatment, they must be treating the observed means as a single phenomena rather than each fund having unique characteristics. Without reading the whole article, I am not sure why they did that.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.