Related Experiment Video
Updated: Jun 26, 2026

The Innovation Arena: A Method for Comparing Innovative Problem-Solving Across Groups
Published on: May 13, 2022
Theory and Inference for Regression Models with Missing Responses and Covariates.
Qingxia Chen1, Joseph G Ibrahim, Ming-Hui Chen
1Qingxia Chen is Assistant Professor, Department of Biostatistics, Vanderbilt University, Nashville, TN 37232, Email: cindy.chen@vanderbilt.edu . Joseph G. Ibrahim is Professor, Department of Biostatistics, University of North Carolina, McGavran-Greenberg Hall, Chapel Hill, NC 27599, Email: ibrahim@bios.unc.edu . Ming-Hui Chen is Professor, Department of Statistics, University of Connecticut, 215 Glenbrook Road, U-4120, Storrs, CT 06269-4120, Email: mhchen@stat.uconn.edu . Pralay Senchaudhuri is Director of Cytel Software Corporation, Cambridge, MA 02139,
This study investigates statistical inference for regression models with missing response and covariate data, assuming Missing at Random (MAR) or Missing Completely at Random (MCAR). It compares three methods: complete case, complete response, and all case analysis.
Area of Science:
- Statistics
- Biostatistics
- Data Science
Background:
- Missing data in regression models pose challenges for accurate statistical inference.
- Existing research often addresses either missing covariates or missing responses, but not both simultaneously.
- Handling both missing response and covariate data requires robust theoretical frameworks.
Purpose of the Study:
- To conduct a theoretical investigation of statistical inference for general regression models with both missing response and covariate data.
- To evaluate and compare three distinct estimation methods: complete case (CC), complete response (CR), and all case (AC) analysis.
- To analyze the theoretical properties, information loss, and asymptotic variances of these methods under MAR and MCAR assumptions.
Main Methods:
- Development of general likelihood expressions for each estimation scenario (CC, CR, AC).
- Application of the Expectation-Maximization (EM) algorithm for estimation scheme development.
- Theoretical analysis within the normal linear model framework, including bias investigation for CC.
- Derivation and comparison of asymptotic variances under MAR and MCAR missing data mechanisms.
Main Results:
- Analytical characterization of information loss for CC, CR, and AC methods.
- Comparison of asymptotic variances across the three methods, highlighting their relative efficiencies.
- Investigation into the potential bias introduced by the complete case analysis method.
- Validation of theoretical findings through a simulation study and a real-world dataset.
Conclusions:
- The study provides a comprehensive theoretical framework for handling missing data in regression models.
- Different missing data handling strategies (CC, CR, AC) exhibit varying levels of efficiency and potential bias.
- The findings offer guidance on selecting appropriate methods for statistical inference in the presence of incomplete response and covariate data.
Related Concept Videos
Multiple Regression
Farmers can use multiple regression to determine the crop yield based on more than one factor, such as water availability, fertilizer, soil properties, etc. Here, the crop yield is the response or dependent variable as it depends on the other independent variables. The analysis requires the construction of a scatter plot...
Assumptions of Survival Analysis
Survival Tree
Building a Survival Tree
Constructing a survival tree begins...
Theory of Attribution I: Correspondent Inference Theory
Regression Analysis
In regression analysis, a regression equation is determined based on the line of best fit– a line that best fits the data points plotted in a graph. This line is also called the regression line. The algebraic equation for the regression line is called the regression equation. It is represented as:
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
