Related Experiment Video
Updated: May 23, 2025

Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
Computational Phenomapping of Randomized Clinical Trial Participants to Enable Assessment of Their Real-World
Phyllis M Thangaraj1, Evangelos K Oikonomou1, Lovedeep S Dhingra1
1Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, New Haven, CT (P.M.T., E.K.O., L.S.D., A.A., R.J., R.K.).
Background:
Assessing the generalizability of randomized clinical trials (RCTs) to real-world patients remains challenging. We propose a multidimensional metric to quantify the representativeness of an RCT cohort in an electronic health record (EHR) population and estimate real-world effects based on individualized treatment effects observed in the RCT.
Methods:
We identified 65 clinical prerandomization characteristics of patients with heart failure with preserved ejection fraction within the TOPCAT (Treatment of Preserved Cardiac Function Heart Failure with an Aldosterone Antagonist Trial) and extracted those features in similar patients in EHR data from 4 hospitals in the Yale New Haven Health System. We then assessed the real-world generalizability of TOPCAT by developing a novel statistic, the phenotypic distance metric, to quantify the representation of TOPCAT participants within EHR patients. Finally, applying a machine learning method to learn individualized treatment effect in TOPCAT participants stratified by region, the United States (US) and Eastern Europe (EE), we predicted spironolactone benefit within the EHR cohorts.
Results:
There were 3445 patients in TOPCAT (median age 69, interquartile range [IQR], 61-76 years, 52% women) and 8121 patients with heart failure with preserved ejection fraction across 4 hospitals (median age range 77, IQR, 68-86; years to 85; IQR, 77-91 years, 54% to 62% women). Across covariates, the EHR patients were more similar to each other than the TOPCAT-US participants (median standardized mean difference 0.065, IQR, 0.011-0.144 versus median standardized mean difference 0.186, IQR, 0.040-0.479). The phenotypic distance metric found a higher generalizability of the TOPCAT-US participants to the EHR patients than the TOPCAT-EE participants. Using a TOPCAT-US-derived model of individualized treatment effect, all EHR patients were predicted to derive benefit from spironolactone treatment, while a TOPCAT-EE-derived model predicted 13% of EHR patients to derive benefit.
Conclusions:
This novel multidimensional metric evaluates the real-world representativeness of RCT participants against corresponding patients in the EHR, enabling the evaluation of an RCT's implication for real-world patients.
Related Concept Videos
Randomized Experiments
Simple randomization
Simple...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Statistical Software for Data Analysis and Clinical Trials
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...

