Related Experiment Video
Updated: Feb 22, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Accounting for misclassification in electronic health records-derived exposures using generalized linear finite
Rebecca A Hubbard1, Eric Johnson1, Jessica Chubak1
1Department of Biostatistics & Epidemiology, University of Pennsylvania, Philadelphia, Pennsylvania (Rebecca A. Hubbard); Group Health Research Institute, Seattle, Washington (Eric Johnson, Jessica Chubak, Karen J. Wernli, Aruna Kamineni); Department of Epidemiology, University of Washington, Seattle, Washington (Jessica Chubak); RAND Corporation, Santa Monica, California (Andy Bogart, Carolyn M. Rutter).
None:
Exposures derived from electronic health records (EHR) may be misclassified, leading to biased estimates of their association with outcomes of interest. An example of this problem arises in the context of cancer screening where test indication, the purpose for which a test was performed, is often unavailable. This poses a challenge to understanding the effectiveness of screening tests because estimates of screening test effectiveness are biased if some diagnostic tests are misclassified as screening. Prediction models have been developed for a variety of exposure variables that can be derived from EHR, but no previous research has investigated appropriate methods for obtaining unbiased association estimates using these predicted probabilities. The full likelihood incorporating information on both the predicted probability of exposure-class membership and the association between the exposure and outcome of interest can be expressed using a finite mixture model. When the regression model of interest is a generalized linear model (GLM), the expectation-maximization algorithm can be used to estimate the parameters using standard software for GLMs. Using simulation studies, we compared the bias and efficiency of this mixture model approach to alternative approaches including multiple imputation and dichotomization of the predicted probabilities to create a proxy for the missing predictor. The mixture model was the only approach that was unbiased across all scenarios investigated. Finally, we explored the performance of these alternatives in a study of colorectal cancer screening with colonoscopy. These findings have broad applicability in studies using EHR data where gold-standard exposures are unavailable and prediction models have been developed for estimating proxies.
More Related Videos
Related Concept Videos
Mechanistic Models: Compartment Models in Individual and Population Analysis
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding in Epidemiological Studies
Statistical Methods for Analyzing Epidemiological Data
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...
Mechanistic Models: Overview of Compartment Models

