FedECA: federated external control arms for causal inference with time-to-event data in distributed settings
Jean Ogier du Terrail1, Quentin Klopfenstein2, Honghao Li2
1Owkin, Inc., New York, NY, USA. jean.duterrail.scientific.contact@gmail.com.
Nature Communications
|August 13, 2025
Summary
Federated learning enables external control arms for drug development without pooling patient data. This method uses inverse probability of treatment weighting for time-to-event outcomes across separate datasets, accelerating clinical research.
Area of Science:
- * Clinical pharmacology and drug development.
- * Biostatistics and real-world evidence generation.
- * Health data privacy and security.
Background:
- * External control arms are crucial for drug development and regulatory approval.
- * Challenges exist in accessing and pooling real-world or historical clinical trial data due to privacy regulations.
- * Centralized data aggregation is often hindered by data protection requirements.
Purpose of the Study:
- * To develop a federated learning method for inverse probability of treatment weighting (IPTW) for time-to-event outcomes.
- * To enable the use of external control arms without centralizing sensitive patient data.
- * To facilitate the comparison of treatment effects across distributed datasets.
Main Methods:
- * Implementation of federated learning to perform IPTW on decentralized patient cohorts.
- * Application of the method in simulated and real-world settings of increasing complexity.
- * Validation using data from three separate patient cohorts for metastatic pancreatic cancer.
Main Results:
- * Demonstrated the feasibility of federated learning for IPTW on time-to-event data across separate cohorts.
- * Successfully applied the method to compare chemotherapy regimens in metastatic pancreatic cancer patients.
- * Showcased the potential for robust comparative effectiveness research without data pooling.
Conclusions:
- * Federated learning offers a viable solution for utilizing external control arms in drug development.
- * The developed method addresses data privacy concerns and facilitates collaborative research.
- * This approach can accelerate the generation of real-world evidence and support regulatory decision-making.
Related Concept Videos
Causality in Epidemiology
834
Causality or causation is a fundamental concept in epidemiology, vital for understanding the relationships between various factors and health outcomes. Despite its importance, there's no single, universally accepted definition of causality within the discipline. Drawing from a systematic review, causality in epidemiology encompasses several definitions, including production, necessary and sufficient, sufficient-component, counterfactual, and probabilistic models. Each has its strengths and...
834
Comparing the Survival Analysis of Two or More Groups
286
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
286
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
126
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
126
Friedman Two-way Analysis of Variance by Ranks
296
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
296
Assumptions of Survival Analysis
197
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
197
Statistical Methods for Analyzing Epidemiological Data
533
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
533


