Related Experiment Video
Updated: Feb 6, 2026

Objectification of Tongue Diagnosis in Traditional Medicine, Data Analysis, and Study Application
Published on: April 14, 2023
Assessing Treatment Effects in Observational Data With Missing Confounders: A Comparative Study of Practical
Brian D Williamson1,2, Chloe Krakauer1, Eric Johnson1
1Biostatistics Division, Kaiser Permanente Washington Health Research Institute, Seattle, Washington, USA.
Doubly-robust methods offer improved efficiency for handling missing data in pharmacoepidemiology research. These under-utilized techniques, generalized raking and targeted maximum likelihood estimation (TMLE), outperform traditional methods like multiple imputation (MI) and inverse-probability weighting (IPW).
Area of Science:
- Pharmacoepidemiology
- Biostatistics
- Health Data Science
Background:
- Pharmacoepidemiology relies on administrative and electronic health records (EHR) for safety and effectiveness studies.
- Missing confounder data is common in these real-world datasets, necessitating robust analytical approaches.
- Multiple imputation (MI) and inverse-probability weighting (IPW) are standard but may not be optimal for missing data.
Purpose of the Study:
- To evaluate and compare the performance of under-utilized doubly-robust methods against traditional MI and IPW for handling missing data in pharmacoepidemiology.
- To provide guidance on selecting appropriate missing data analysis methods based on performance across various scenarios.
Main Methods:
- Investigated two doubly-robust estimators: generalized raking and inverse probability-weighted targeted maximum likelihood estimation (TMLE).
- Conducted extensive numerical studies with synthetic data across diverse missingness and data-generating scenarios, including rare outcomes and high missingness proportions.
- Utilized plasmode simulation studies emulating large EHR cohort data to assess performance in a rare-outcome setting with substantial missing confounder data (>50%).
Main Results:
- Doubly-robust methods demonstrated superior performance, particularly in terms of the bias-variance trade-off, compared to MI and IPW across various scenarios.
- Generalized raking and TMLE showed greater efficiency and robustness, especially under conditions of high missingness and rare outcomes.
- The study identified specific scenarios where doubly-robust methods significantly reduce bias and improve precision in effect estimation.
Conclusions:
- Doubly-robust methods, specifically generalized raking and TMLE, represent valuable, yet under-utilized, tools for missing data analysis in pharmacoepidemiology.
- These methods offer improved statistical efficiency and robustness, leading to more reliable safety and effectiveness estimates from real-world data.
- Researchers are encouraged to adopt doubly-robust approaches for handling missing confounder data to enhance the validity of pharmacoepidemiological findings.
Related Concept Videos
Data Collection by Observations
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
Assessment of the Gastrointestinal System I: Subjective Data
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
How Data are Classified: Categorical Data
Data are classified based on whether they are measurable or not. Categorical data cannot be measured; instead, it can be divided into categories. For example, if Y denotes a person's party affiliation, some examples of Y include...
Confounding in Epidemiological Studies
How Data are Classified: Numerical Data
Quantitative data may be either discrete or continuous. All quantitative data that take on only specific numerical...

