Related Experiment Video
Updated: Sep 28, 2026

Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
Optimal subsampling methods for large clustered imbalanced data based on generalized estimating equations
Tingting Yu1, Sehwan Kim1, Rui Wang1,2
1Department of Population Medicine, Harvard Pilgrim Healthcare Institute and Harvard Medical School, Boston, MA 02215, United States.
Abstract:
Generalized estimating equations (GEEs) are widely used for modeling clustered data while accounting for within-cluster correlations. However, fitting GEEs on massive datasets can be computationally prohibitive. Subsampling offers a practical solution, but efficient subsampling strategies for large clustered data, particularly with rare events, remain underexplored. Simple random subsampling can lead to a substantial loss of rare-event information, resulting in inefficient estimation, computational instability, and increased finite-sample bias. Additionally, intra-cluster correlation complicates inference on subsampled data, requiring methods that preserve correlation structure while adjusting for unequal sampling probabilities. To address these challenges, we develop an individual-level sampling scheme and derive optimal subsampling probabilities based on GEEs. We propose a subsampling-based weighted GEE (SaWGEE) estimator that efficiently approximates the full-data GEE estimator and establish its asymptotic properties as the number or size of clusters grows. The optimal subsampling probabilities minimize the asymptotic mean-square-error or its upper bound of the SaWGEE estimator, conditional on observed data, under an independent or exchangeable working correlation structure. Simulation studies evaluate the performance of the SaWGEE estimator against conventional subsampling strategies, such as random or 1:1 case-control sampling, and the standard full-data GEEs when feasible. We apply the proposed method to the 2022 CDC survey data, where observations are clustered by state, to construct a risk-adjustment model for heart attack incidence.
Related Concept Videos
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Choosing Between z and t Distribution
Stratified Sampling Method
To choose a stratified sample, divide the population into groups called strata and then take a...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...
One-Way ANOVA: Unequal Sample Sizes