Skip to content
All library documents

Assessing Missing Covariate Data Before Excluding Regression Variables

Article Quant Q&A · Author: Nguyen Lis

Summary

The document raises a regression-design question: whether to exclude a covariate when it has substantially fewer observations than the other variables. The author gives an example in which one variable has 80,000 observations while others have 100,000, and asks for a reference or rationale for dropping it. They suspect that reduced sample size, within-sample variation, and explanatory power may matter, but provide no analysis or proposed rule.

Because the text contains only the question, it offers no evidence that a particular missingness threshold justifies exclusion. The issue is left unresolved: deciding what to do would require examining why values are missing, how complete-case deletion changes the estimation sample, and whether imputation or another missing-data method is appropriate. The example therefore frames a statistical modeling concern rather than teaching a tested procedure or giving a citation.

Key ideas

  • The document asks whether a covariate should be excluded when it has fewer observed values than other variables.
  • It illustrates the question with a covariate observed less often than the rest of the data.
  • Dropping a variable may also shrink the sample used for estimation.
  • The text does not establish a missingness threshold or provide a reference supporting exclusion.
  • The choice of method remains unresolved in the document.

Tags

Full text
# What is the reference for excluding a covariate if there is many missing observation?


# What is the reference for excluding a covariate if there is many missing observation?












Normally, I exclude a covariate out of the regression equation if there are many missing observations, let's say it is one-fifth less observation compared to other variables' observations in general (saying 80,000 compared to 100,000).

But now, when reflecting back, I am wondering if there is any reference or explanation for excluding action like that (excluding a covariate having many missing observation)? I think it may relate to the within-sample standard variation and explanation power due to the sample shrinking but I am not sure about that.

Shown in full with attribution under the source's licence. Licence: CC BY-SA 4.0 (Stack Exchange)

This summary was written by Stratmill's research agent from the original; it is not a copy of the source.