Related Experiment Videos
A Nonparametric Data-Fusion Approach for Identification and Estimation of Nonignorable Missing Data With Shadow
1School of Data Science, Fudan University, Shanghai, China.
Abstract:
Missing data can pose fundamental challenges to statistical inference, with nonignorable missing or missing not at random (MNAR) presenting the most severe methodological difficulties. Despite substantial advances in MNAR inference methods, several limitations remain. These generally include interpretability issues, non-identifiability, unverifiable parametric assumptions, and reliance on external information for which clear practical guidance is often lacking-particularly in sensitivity and Bayesian methods. In this paper, motivated by a real-world COVID-19 dataset from the Centers for Disease Control and Prevention (CDC), we study an MNAR scenario in the presence of a shadow variable that may itself also be subject to MNAR. A shadow variable is associated with the primary variable that is MNAR, but is conditionally independent of that variable's missingness given the primary variable and other covariates. Under the pattern mixture model framework, we propose a nonparametric inference method that can leverage external data to explicitly address the identification and estimation problems in the primary data. The method requires inference only of an observed data density and an odds of missing function, with the latter forming the core of our approach. We further provide two multiple-imputation-based estimation strategies, enhancing transparency and interpretability. The proposed framework can accommodate both covariate MNAR and outcome MNAR settings, as well as different variable types, and extends the existing shadow variable paradigm for MNAR data by relaxing the common assumption that the shadow variable must be fully observed. We evaluate the proposed method through simulation studies and apply it to CDC COVID-19 surveillance data.
Related Concept Videos
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
On...
Introduction to Nonparametric Statistics
One of...
Friedman Two-way Analysis of Variance by Ranks
Estimating Population Mean with Unknown Standard Deviation
William S. Gosset (1876–1937) of the Guinness...
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...