Related Experiment Video
Updated: Oct 1, 2025

Development of an Individual-Tree Basal Area Increment Model using a Linear Mixed-Effects Approach
Published on: July 3, 2020
Semiparametric marginal methods for clustered data adjusting for informative cluster size with nonignorable zeros
Biyi Shen1, Chixiang Chen2, Vernon M Chinchilli1
1Division of Biostatistics and Bioinformatics, Department of Public Health Sciences, Penn State College of Medicine, Hershey, PA, USA.
This study introduces novel methods, weighted within-cluster resampling (WWCR) and dual-weighted generalized estimating equations (WWGEE), to effectively analyze clustered data, including informative zero cluster sizes. These approaches improve data utilization in clinical and observational studies.
Area of Science:
- Biostatistics
- Clinical Epidemiology
- Longitudinal Data Analysis
Background:
- Clustered and longitudinal data are common in clinical research, often collected via real-time monitoring.
- Cluster size can be informative, potentially correlating with disease status.
- Existing methods often exclude nonignorable zero-sized clusters, losing valuable data.
Purpose of the Study:
- To develop statistical methods that incorporate informative nonignorable zero-sized clusters.
- To improve the analysis of clustered data by utilizing information from both event-free and event-occurring participants.
- To address limitations in current statistical approaches for longitudinal and clustered health data.
Main Methods:
- Proposed weighted within-cluster resampling (WWCR) method.
- Introduced asymptotically equivalent dual-weighted generalized estimating equations (WWGEE).
- Employed inverse probability weighting techniques to handle informative cluster sizes.
Main Results:
- Theoretically established asymptotic properties of the proposed methods.
- Demonstrated the finite-sample behavior through extensive simulations.
- Showcased advantageous performance compared to existing approaches in the ASSESS-AKI study.
Conclusions:
- The WWCR and WWGEE methods effectively utilize data from informative zero-sized clusters.
- These novel methods offer improved statistical power and accuracy in analyzing clustered health data.
- The findings have significant implications for clinical trials and observational studies with complex data structures.
More Related Videos
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
08:27Applying an eMASS Customization Program as a Research Tool to Evaluate Consumer Benefits
Published on: September 27, 2019
Related Concept Videos
Cluster Sampling Method
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
Parametric Survival Analysis: Weibull and Exponential Methods
Weibull Distribution
The Weibull distribution is a flexible model used in parametric survival analysis. It can handle both increasing and decreasing hazard rates, depending on its shape parameter...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the...
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
One-Way ANOVA: Equal Sample Sizes
Different sample means can result in different values for the variance estimate: variance between samples. This is because the variance between samples is calculated as the product of the sample size and the variance between the...