Resampled Efficient Frontiers and Simulation Sample Length
Summary
The document considers a resampled efficient frontier procedure that simulates multivariate normal monthly returns. An investment horizon of six years is described, while the lecturer’s program generates twelve years of data and uses a six-year slice to estimate expected returns and covariance for the procedure. The question is whether generating the longer series changes the result when only six years are used.
The accepted response says that the unused observations do not affect the estimates from the selected slice, while a longer simulated sample can produce estimates closer to the assumed population parameters when those observations are used. It suggests that holding back the second half may be intended for an out-of-sample assessment of the resampled method, whose value is framed as out-of-sample performance rather than in-sample superiority to mean-variance optimization. The explanation is tentative: the code may instead contain an error, and the document gives no empirical comparison or details about the precise sampling implementation.
Key ideas
- Estimates based on a six-year slice use only the observations in that slice.
- A longer simulated sample can improve estimates of the assumed population parameters if the extra data are included.
- A held-back portion of simulated returns could support an out-of-sample evaluation.
- The document presents the holdout explanation as a possibility rather than a confirmed account of the code.
Tags
Full text
# Resampled efficient frontier length of simulation # Resampled efficient frontier length of simulation I was provided with a VBA program from my lecturer that applies the resampled efficient frontier. We have an investment horizon $T$ (6 years) and he uses a multivariate normal distribution with parameters $\mu$ and $\Sigma$ and the code simulates $12$ years of monthly data. He does this a few hundred times, but only takes the top half of the returns matrix, $R$, (i.e. takes a $6$ year slice of it) to calculate the estimates of $\mu$ and $\Sigma$ in the REF procedure. Does it make any difference whether we simulate $6$ years or $12$ years under MVN if we're just going to be slicing off $6$ years of simulated returns from the top of the simulated $R$ no matter what? ## Answer by vanguard2k (score 0, accepted) https://quant.stackexchange.com/a/4231 It doesnt make any difference. Of course with 12 years under MVN your estimates will be closer to the "true" values of the population. Maybe you want to use the second part of $R$ to perform an out of sample test of the REF to show that the strength of the REF lies not in its in-sample performance (where it must be inferior to MV optimization) but rather in its out-of-sample performance but we can only guess on this part. It could also be an error in the code but I think to hold half of the data back for the out-of-sample test should be the reason.
Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)
This summary was written by Stratmill's research agent from the original; it is not a copy of the source.