Related Experiment Video
Updated: Jan 11, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluation of Methods Adjusting for Unmeasured Confounding Using Large Healthcare Databases: An Empirical Study
Chi-Hong Duong1, Sylvie Escolano1, Romain Demailly1,2
1High-Dimensional Biostatistics for Drug Safety and Genomics, CESP, Université Paris-Saclay, UVSQ, Université Paris-Sud, Inserm, Villejuif, France.
Abstract:
With the growing availability of large healthcare databases for clinical science, mitigating unmeasured confounding has emerged as a major issue in pharmacoepidemiologic studies. Extensions of causal inference methods to high-dimensional settings could help address this problem, but studies comparing their performance in real-world databases are still lacking. This study aims to compare the ability to reduce the measured and indirectly measured confounding of three causal inference methods adapted to a real-world high-dimensional database using a machine learning LASSO algorithm: G-computation (GC), Targeted Maximum Likelihood estimation (TMLE) and Propensity Score with overlap or stabilized inverse probability treatment weighting. This large-scale empirical study was based on the French National Healthcare Claims Database (SNDS), consisting of 2,172,702 pregnancies 22 weeks of gestation over the period 2011-2014. We used a set of 42 negative and 13 positive reference drugs related to prematurity risk. For each reference drug, the logarithm of the odds ratio for prematurity and its 95% confidence interval were estimated using each method. The proportions of false positive and true positive associations were calculated and compared between the methods. All methods yielded fewer false positives than a crude model based on a minimal set of adjusted covariates. TMLE produced the lowest proportion of false positives (45.2%), followed by GC (47.6%). GC yielded the highest proportion of true positives (92.3%). Our results confirm the interest of causal inference methods exploiting the wealth of data in healthcare databases, especially GC in terms of performance and ease of implementation.
More Related Videos
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Factors Affecting Drug Response: Overview
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Regression Toward the Mean
Confounding in Epidemiological Studies

