Related Experiment Video
Updated: Aug 5, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Sample Size Reduction by Applying ML Based Causal Inference Methods
Ahrim Youn1, Gang Han1, Emma Gerard2
1Evidence Generation and Decision Science, Sanofi, Morristown, New Jersey, USA.
Abstract:
We conducted a comprehensive comparative analysis of causal machine learning (ML) methods to assess their utility in improving the efficiency of clinical trial designs with or without historical data. Specifically, we compared standard ANCOVA analysis used in a Randomized Controlled Trial (RCT) with several causal ML methods, including PROCOVA, TMLE, DML, and GRF. PROCOVA is gaining popularity in RCT design and requires a historical data for prognostic scores, but other methods can be applied with or without such data. Our primary focus was on strict RCT setting without borrowing historical control data, though we also explored the impact of borrowing data. The historical data used consisted of placebo data from two Phase 3 Ophthalmology studies with a continuous primary endpoint. We employed a generative AI approach, specifically Generative Adversarial Networks (GANs), to simulate RCT data from the historical data under various scenarios, varying treatment effects with and without treatment effect heterogeneity, RCT sizes, and bias. Results showed that causal ML methods can increase power even without borrowing historical data. For example, TMLE increased effective sample size by 21% in one scenario. In scenarios with borrowing of controls, PROCOVA increased power while controlling type 1 error, showing robustness to model misspecification.
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Censoring Survival Data
Methods of Medium Optimization
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
