Related Experiment Video
Updated: May 2, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Measuring the Impact of Data Quality and Computable Phenotypes on Potential Racial Disparities in Predicting
Priyanka D Sood1, Star Liu2, Chintan Pandya1
1Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, USA.
Introduction:
Type 2 diabetes (T2D) computable phenotypes identify different denominator populations for downstream tasks. Differences in racial composition could introduce bias and lead to disparate disease management. The objective of this study was to assess potential racial disparities in predicting T2D healthcare utilization introduced by data quality and computable phenotypes.
Methods:
Four published and one local T2D phenotypes were applied to the EHR and claims datasets of a large academic medical center. Population characteristics were compared across phenotypes, stratified by race. We induced data incompleteness, inaccuracy, and untimeliness to measure the impact on denominator racial composition. We trained logistic classification models on each of the phenotype-specific populations separately and compared disparities in utilization prediction (i.e., inpatients (IP) and emergency room (ER) admissions). Model performance, such as mean AUC and positive/negative predictive values, were compared across phenotypes, stratified by race.
Results:
Different T2D computable phenotypes identified populations with modestly different racial compositions. Black T2D patients had the highest average admissions to ER compared to other racial groups. Induced data quality challenges diminished patient counts across all racial groups proportionally. Charlson comorbidity score had the highest odds ratio in predicting IP and ER admissions across phenotypes and race groups. Specific T2D phenotypes showed the highest and lowest mean AUCs in predicting IP and ER admissions in Black and White populations; however, such results were not observed among Asian/Other populations.
Conclusion:
Utilization prediction differed among phenotypes and race groups. Understanding the complexities behind phenotypes, data quality, and predictive models could mitigate health disparity further downstream and inform clinical research and disease management.
Related Concept Videos
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Type II Diabetes I: Introduction
Analysis of Population Pharmacokinetic Data
Diabetes Mellitus: Type 2 and Gestational
Pharmacogenomics: Identification of New Drug Targets
Regression Toward the Mean
