Related Experiment Video
Updated: Jan 26, 2026

Electrochemical Impedance Spectroscopy as a Tool for Electrochemical Rate Constant Estimation
Published on: October 10, 2018
A Note on the Effect of Data Clustering on the Multiple-Imputation Variance Estimator: A Theoretical Addendum to ,
Yulei He1, Iris Shimizu1, Susan Schappert1
1National Center for Health Statistics, Centers for Disease Control and Prevention, Hyattsville, MD, 20782, U.S.A.
Abstract:
Multiple imputation is a popular approach to handling missing data. Although it was originally motivated by survey nonresponse problems, it has been readily applied to other data settings. However, its general behavior still remains unclear when applied to survey data with complex sample designs, including clustering. Recently, Lewis et al. (2014) compared single- and multiple-imputation analyses for certain incomplete variables in the 2008 National Ambulatory Medicare Care Survey, which has a nationally representative, multistage, and clustered sampling design. Their study results suggested that the increase of the variance estimate due to multiple imputation compared with single imputation largely disappears for estimates with large design effects. We complement their empirical research by providing some theoretical reasoning. We consider data sampled from an equally weighted, single-stage cluster design and characterize the process using a balanced, one-way normal random-effects model. Assuming that the missingness is completely at random, we derive analytic expressions for the within- and between-multiple-imputation variance estimators for the mean estimator, and thus conveniently reveal the impact of design effects on these variance estimators. We propose approximations for the fraction of missing information in clustered samples, extending previous results for simple random samples. We discuss some generalizations of this research and its practical implications for data release by statistical agencies.
More Related Videos
06:01Visualization and Quantification of High-Dimensional Cytometry Data using Cytofast and the Upstream Clustering Methods FlowSOM and Cytosplore
Published on: December 12, 2019
10:46A Method of Trigonometric Modelling of Seasonal Variation Demonstrated with Multiple Sclerosis Relapse Data
Published on: December 9, 2015
Related Concept Videos
Variance
The standard deviation measures the spread in the same units as the data....
Theoretical Approaches to Psychological Disorder
Biological approach
The biological approach posits that internal, organic factors are the primary causes of such disorders. This perspective emphasizes brain structure and function, genetic predispositions, and neurotransmitter imbalances. For example, schizophrenia has been associated with both genetic...
Theoretical Foundations of Nursing Practice
Theories provide a perspective to assess patients' conditions and organize data and methods. They also assist in analyzing and interpreting information. They represent a...
Friedman Two-way Analysis of Variance by Ranks
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...