Related Experiment Video
Updated: Dec 17, 2025

Author Spotlight: Integrated Multi-Omics Analysis for Unveiling Multicellular Immune Signatures in Clinical Heart Attack Cohorts
Published on: September 20, 2024
Impact of Diverse Data Sources on Computational Phenotyping
Liwei Wang1, Janet E Olson2,3, Suzette J Bielinski2
1Division of Digital Health Sciences, Department of Health Sciences Research, Mayo Clinic, Rochester, MN, United States.
Integrating diverse data sources significantly improves computational phenotyping accuracy for rheumatoid arthritis and type 2 diabetes mellitus, overcoming limitations of data fragmentation in electronic health records.
Area of Science:
- Biomedical Informatics
- Health Informatics
- Computational Biology
Background:
- Electronic health records (EHRs) offer extensive data for computational phenotyping.
- Data fragmentation within single EHR sources limits the accuracy of automated phenotype extraction.
- Existing computational phenotyping algorithms face challenges due to incomplete data.
Purpose of the Study:
- To investigate the impact of data fragmentation on computational phenotyping algorithms.
- To evaluate the performance of rheumatoid arthritis (RA) and type 2 diabetes mellitus (T2DM) phenotyping algorithms using diverse data sources.
- To compare algorithm performance between a single health system (Mayo EHRs) and a multi-system linked dataset (Rochester Epidemiology Project - REP).
Main Methods:
- Utilized two established computational phenotyping algorithms for RA and T2DM.
- Applied algorithms to single-source Mayo EHR data and multi-source REP data.
- Assessed algorithm performance using metrics such as positive predictive value (PPV) and false-negative rate (FNR).
Main Results:
- Data fragmentation markedly impacted both RA and T2DM case selection accuracy.
- REP data demonstrated superior performance with higher PPV (97.2% for RA, 98.3% for T2DM) and lower FNR (5.2% for RA, 3.3% for T2DM) compared to Mayo EHRs.
- Mayo EHRs showed lower PPV (91.4% for RA, 92.4% for T2DM) and higher FNR (26.6% for RA, 14% for T2DM).
- Biases were also observed in T2DM control selection from Mayo data (PPV 91.2%, FNR 1.2%).
Conclusions:
- Diverse data sources, like the REP, significantly enhance the accuracy of computational phenotyping by mitigating data fragmentation.
- Utilizing multi-system linked EHR data is crucial for reliable phenotype extraction.
- Addressing data fragmentation is essential for advancing phenotype-driven research and improving healthcare delivery through EHR data.
Related Concept Videos
Genomics
Analysis of Population Pharmacokinetic Data
Background and Environment Affect Phenotype
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
Evolutionary Relationships through Genome Comparisons
Genetic Screens
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Incomplete Dominance

