Related Experiment Video
Updated: Jan 5, 2026

15:30
A Telemetric, Gravimetric Platform for Real-Time Physiological Phenotyping of Plant–Environment Interactions
Published on: August 5, 2020
12.4K
High-throughput multimodal automated phenotyping (MAP) with application to PheWAS.
Katherine P Liao1,2,3, Jiehuan Sun3,4, Tianrun A Cai1,2,3
1Division of Rheumatology, Immunology, and Allergy, Brigham and Women's Hospital, Boston, MA, USA.
Journal of the American Medical Informatics Association : JAMIA
|October 16, 2019
Summary
This study introduces a Multimodal Automated Phenotyping (MAP) algorithm that accurately identifies patient phenotypes by combining electronic health records (EHR) and natural language processing (NLP). The MAP algorithm enhances phenotyping accuracy and scalability for large-scale translational research.
Area of Science:
- Biomedical Informatics
- Translational Research
- Computational Biology
Background:
- Accurate and efficient patient phenotyping is crucial for translational studies but remains a significant bottleneck.
- Electronic health records (EHR) linked with biorepositories offer a powerful platform for research.
- Current phenotyping methods often rely on International Classification of Diseases (ICD) codes, which may lack granularity.
Purpose of the Study:
- To develop an automated, high-throughput phenotyping method.
- To integrate International Classification of Diseases (ICD) codes and narrative data extracted using natural language processing (NLP).
- To improve the accuracy and efficiency of patient phenotyping for large-scale studies.
Main Methods:
- Developed a mapping method using the Unified Medical Language System to identify relevant ICD and NLP concepts.
- Employed latent mixture models to jointly analyze aggregated ICD and NLP counts alongside healthcare utilization.
- Created the Multimodal Automated Phenotyping (MAP) algorithm to predict phenotype probability and classify patients.
Main Results:
- The MAP algorithm demonstrated superior or comparable performance (AUC, F-score) to ICD codes alone across 16 phenotypes.
- Automated feature extraction achieved accuracy comparable to manual curation (AUCMAP 0.943 vs. AUCmanual 0.941).
- Phenome-wide association studies (PheWAS) using MAP showed higher power in detecting validated associations compared to ICD-based methods.
Conclusions:
- The MAP approach significantly enhances phenotype definition accuracy and scalability.
- This automated method facilitates large-scale phenotyping essential for studies like PheWAS.
- MAP represents a significant advancement in leveraging EHR and NLP for translational research.

