Skip to content
All library documents

Bootstrap Sampling for Confidence Intervals with Correlated Data

Article Quant Q&A · Author: Ksnapp

Summary

This document considers confidence intervals for a nonlinear function of three random variables, including a skewed variable and a correlated pair. It compares resampling each variable independently with resampling complete observed rows. Because row-wise resampling retains the observed association between variables, the questioner favors that method when estimating the sampling distribution of the function.

The question also asks whether the standard deviation of bootstrap estimates can be used to form confidence bounds. The supplied reply challenges the problem’s assumptions, arguing that the variables, as described, are independent by nature and that confidence intervals may therefore be irrelevant. That response does not explain how to reconcile this claim with the stated sample correlation, nor does it give a bootstrap interval procedure. The material is thus useful as a prompt to define the target of inference and dependence structure carefully, but it is not a complete guide to confidence intervals for skewed or unbounded statistics.

Key ideas

  • Independent resampling can discard dependence present in paired observations.
  • Resampling complete rows preserves the observed joint relationships among variables.
  • The target estimate and assumptions about dependence must be clear before constructing an interval.
  • The response questions whether confidence intervals are appropriate but does not provide a complete method.

Tags

Full text
# How would I develop confidence bounds for a function of 3 random variables, 2 of which are correlated?


# How would I develop confidence bounds for a function of 3 random variables, 2 of which are correlated?












I am tasked with developing confidence intervals for the function x = 1 - |(a+b)/c| where a, b and c are random variables. a and b are normally distributed, but c is heavily skewed left. further there b and c are correlated (Pearson = .21). x belongs to [-inf, 1] and is skewed left because of its domain. I have 2000 "rows" of data.

I can't see an analytic solution for this equation/problem. I was thinking this would be a great opportunity to bootstrap an estimate. I've approached it two ways; 1) resample a, b, and c independently. Then, calculate x^ each time, then take summary statistics for every iteration. Iterate 10,000 or so times. 2) resample the rows of data for a, b, and c. Calculate x^ and take summary statistics. Again, iterate 10,000 or so times.

I think the correct approach is 2 because then the correlation between b and c will be present in the data. Please tell me if I'm right.

Now once if have a big collection of means of x^ can I use the standard deviation of the means of x^ to get the 95% or whatever % confidence bounds?

## Answer by Chris (score 1)

https://quant.stackexchange.com/a/47256

Confidence intervals are applied to estimates to give a sense for potential error. If a, b, and c are R.V. as you described, they're independent by nature, making confidence intervals irrelevant. Either you're misunderstanding what was asked of you or your problem is misspecified (or you just misspecified it here).

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.