A one-shot, lossless algorithm for cross-cohort learning in mixed-outcomes analysis.
Ruowang Li1, Luke Benz2, Rui Duan2
1Department of Computational Biomedicine, Cedars-Sinai Medical Center, Los Angeles, CA 90049, USA.
Patterns (New York, N.Y.)
|October 3, 2025
Summary
A new algorithm, mixWAS, efficiently integrates distributed electronic health record (EHR) data using summary statistics. This lossless method enhances cross-cohort genetic association discovery for improved healthcare research insights.
Area of Science:
- Bioinformatics
- Computational Biology
- Genetics
Background:
- Integrating diverse electronic health record (EHR) datasets for cross-cohort studies is challenging due to data heterogeneity, distributed storage, and privacy issues.
- Traditional data pooling or harmonization methods can be inefficient and limit the scope of cross-cohort analyses.
- Existing approaches may not preserve cohort-specific information or support simultaneous multi-outcome analyses.
Purpose of the Study:
- To introduce mixWAS, a novel one-shot, lossless algorithm for efficient integration of distributed EHR datasets.
- To enable precise cross-cohort learning while preserving cohort-specific covariate associations and supporting mixed-outcome analyses.
- To improve the accuracy and efficiency of identifying genetic associations across multiple EHR datasets.
Main Methods:
- mixWAS algorithm utilizes summary statistics for lossless integration of distributed EHR data.
- The algorithm preserves cohort-specific covariate associations.
- Simulations were conducted to compare mixWAS performance against conventional methods.
Main Results:
- mixWAS demonstrated superior accuracy and efficiency compared to traditional methods in simulations.
- Application to seven US EHR cohorts identified 4,530 significant cross-cohort genetic associations for traits including blood lipids, BMI, and circulatory diseases.
- Validation using an independent UK EHR dataset confirmed 97.7% of the identified associations, highlighting the algorithm's robustness.
Conclusions:
- mixWAS enables lossless integration of distributed EHR data, enhancing precision in multi-outcome analyses.
- The algorithm facilitates robust cross-cohort genetic association discovery, improving healthcare research.
- mixWAS expands the potential for generating actionable insights from large-scale, distributed health data.
Related Concept Videos
Comparing the Survival Analysis of Two or More Groups
553
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
553
Crossover Experiments
4.5K
Crossover experiments, also called the repeated-measurements design, is a study design in which all experimental units are exposed to all treatments in different periods. Crossover experiments are generally used in psychology, the pharmaceutical industry, agriculture, and medicine.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
Crossover designs are performed even with smaller sample sizes since the samples can act as their controls. These are better than simple randomized trials since patients are exposed to all the treatments.
4.5K
Cross-Sectional Research
12.4K
In cross-sectional research, a researcher compares multiple segments of the population at the same time. If they were interested in people's dietary habits, the researcher might directly compare different groups of people by age. Instead of following a group of people for 20 years to see how their dietary habits changed from decade to decade, the researcher would study a group of 20-year-old individuals and compare them to a group of 30-year-old individuals and a group of 40-year-old...
12.4K
Assumptions of Survival Analysis
392
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
392
Longitudinal Studies
477
Longitudinal studies are also widely used in other medical and social science fields. For instance, in cardiovascular research, they can monitor patients' health over decades to identify risk factors for heart disease, such as high cholesterol or smoking, and evaluate the long-term effectiveness of preventive measures. Similarly, in mental health studies, researchers might follow individuals from adolescence into adulthood to understand the development and progression of conditions like...
477
Truncation in Survival Analysis
577
Truncation in survival analysis refers to the exclusion of individuals or events from the dataset based on specific criteria related to the time of the event. This exclusion can happen in two primary forms: left truncation and right truncation.
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
Left truncation occurs when individuals who experienced the event of interest before a certain time are not included in the study. This is often due to a "delayed entry" into the study where only those who survive until a certain entry point are...
577


