Related Experiment Video
Updated: Jan 11, 2026

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Evaluation of Methods Adjusting for Unmeasured Confounding Using Large Healthcare Databases: An Empirical Study
Chi-Hong Duong1, Sylvie Escolano1, Romain Demailly1,2
1High-Dimensional Biostatistics for Drug Safety and Genomics, CESP, Université Paris-Saclay, UVSQ, Université Paris-Sud, Inserm, Villejuif, France.
Causal inference methods like G-computation (GC) and Targeted Maximum Likelihood estimation (TMLE) effectively reduce confounding in large healthcare databases. GC demonstrated superior performance in identifying true positive drug-prematurity associations.
Area of Science:
- Pharmacoepidemiology
- Causal Inference
- Health Data Science
Background:
- Large healthcare databases offer opportunities for pharmacoepidemiologic research.
- Mitigating unmeasured confounding is crucial for valid causal inference in these datasets.
- Comparing causal inference methods in real-world, high-dimensional data is essential.
Purpose of the Study:
- To compare the performance of G-computation (GC), Targeted Maximum Likelihood estimation (TMLE), and Propensity Score methods in reducing measured and unmeasured confounding.
- To evaluate these methods using a machine learning LASSO algorithm on a high-dimensional, real-world healthcare database.
- To assess the ability of these methods to identify true and false positive drug-prematurity associations.
Main Methods:
- Utilized the French National Healthcare Claims Database (SNDS) with 2,172,702 pregnancies (≥22 weeks gestation, 2011-2014).
- Employed G-computation (GC), Targeted Maximum Likelihood estimation (TMLE), and Propensity Score methods adapted for high-dimensional data.
- Assessed 42 negative and 13 positive reference drugs for prematurity risk, estimating odds ratios and confidence intervals.
Main Results:
- All causal inference methods outperformed a crude model in reducing false positives.
- Targeted Maximum Likelihood estimation (TMLE) yielded the lowest false positive rate (45.2%).
- G-computation (GC) achieved the highest true positive rate (92.3%) and demonstrated good performance and ease of implementation.
Conclusions:
- Causal inference methods are valuable for leveraging large healthcare databases.
- G-computation (GC) shows particular promise for pharmacoepidemiologic studies due to its performance and implementation simplicity.
- These findings support the use of advanced causal inference techniques to address confounding in real-world health data research.
More Related Videos
Related Concept Videos
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Factors Affecting Drug Response: Overview
Pharmacovigilance
This process, termed pharmacovigilance, aims to detect, evaluate, and minimize harmful effects related to medication use. The data collection for pharmacovigilance depends on spontaneous reporting systems, where healthcare professionals or patients voluntarily report suspected ADRs.
In some cases, there...
Regression Toward the Mean
Confounding in Epidemiological Studies

