Skip to content
All library documents

Using Quasi-Monte Carlo for Regression Simulation

Article Quant Q&A · Author: Bazman

Summary

The document considers whether quasi Monte Carlo (QMC) can improve a simulation that repeatedly generates regression errors and compares estimators by mean squared error. It contrasts pseudorandom sampling, where a standard error estimate can support a precision-based stopping rule, with low-discrepancy sequences, whose error cannot be estimated from the sample standard deviation in the same way. QMC users may instead choose the sample count in advance or monitor whether the running estimate has stabilized.

For comparing Halton, Sobol, and other sequences, the answer recommends evaluating a suite of representative problems and comparing computation time at a given precision. It cites a historical equity derivatives comparison in which Niederreiter performed slightly better than Sobol and Faure. That result is specific to the tested problems; QMC is not presented as universally superior, and the document gives no direct benchmark for the regression simulation in question.

Key ideas

  • QMC may improve convergence for some simulations, but its benefit depends on the problem being estimated.
  • Pseudorandom sampling supports standard error based stopping rules that do not transfer directly to low-discrepancy sequences.
  • QMC runs can use a predetermined sample count or stop when the running estimate changes very little.
  • Compare candidate sequences across representative problems using computation time at a given precision.
  • A historical equity derivatives comparison favored Niederreiter only by a small margin over Sobol and Faure.

Tags

Full text
# Quasi Monte Carlo in Matlab


# Quasi Monte Carlo in Matlab












I want to use Quasi Monte Carlo to try and improve the convergence of a simulation I am running.

The random numbers are simply to produce the observation errors for a standard linear regression model. Which is then estimated using a number of different regression techniques. This is done repeatedly to estimate the mean square error of each model.

I'm fairly new to Quasi Monte Carlo but is is likely to help in this situation I am just using it to produce 10k random numbers. It seems that generally I can expect quicker convergence of the order of (1/n) rather than n^(-0.5):

http://en.wikipedia.org/wiki/Quasi-Monte_Carlo_method

However it also states that the QMC numbers are not truly random, so I just wonder what the implications might be for any statistical tests I might want to run on the results.

1.) I guess what I want to know are the pros and cons of MC v QMC. (would you always want to use QMC if its available?) 2.) What tests can I use to ascertain which is best for my application? (seems any test that depends on the numbers being truly random will fail?)

I know that this can be done in Matlab using

q = qrandstream('halton',NSteps,'Skip',1e3,'Leap',1e2); RandMat = qrand(q,NRepl); z_RandMat = norminv(RandMat,0,1);

which is taken from this paper.

http://papers.ssrn.com/sol3/cf_dev/AbsByAuth.cfm?per_id=1519239

It seems there are other low discrepancy numbers such as Sobol sequence available in Matlab and again would just like to know what tests I can use to ascertain which is best for my situation.

## Answer by Brian B (score 2)

https://quant.stackexchange.com/a/10043

One of the main things you give up is a simple halting condition for your estimation algorithm. With pseudorandom numbers, the algorithm can keep track of the standard error, and stop when it has passed a threshold:

```
error_est = Inf
n = 0
while not error_est < target_precision:
    n = n + 1
    x = new_random_sample()
    samples.append( F(x) )
    error_est = 3 * std_deviation( samples ) / sqrt(n)
value_est = mean( samples )
```

Since with QR sequences your error estimate can't be taken from the standard deviation, you cannot apply this algorithm. Instead, practitioners often either just choose a number of samples to take a priori, or set up a halting condition that checks for the running mean to be changing very little.

As for tests of one sequence versus another, the typical approach is to set up a large suite of sample problems. The winner is then the QR sequence that has the best computation time to precision tradeoff on the problem suite.

When we did this for equity exotics in the 1990s, we found Niederreiter sequences beat Sobol and Faure, though by a very small amount.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.