Related Experiment Video
Updated: Jan 5, 2026

A Telemetric, Gravimetric Platform for Real-Time Physiological Phenotyping of Plant–Environment Interactions
Published on: August 5, 2020
Ensembles of natural language processing systems for portable phenotyping solutions
Cong Liu1, Casey N Ta1, James R Rogers1
1Department of Biomedical Informatics, Columbia University, New York, NY 10032, USA.
Ensemble natural language processing (NLP) methods improve automated phenotype extraction from electronic health records (EHRs), enhancing portability across diverse patient cohorts. This approach offers a more reproducible and efficient solution for clinical phenotyping compared to individual NLP systems.
Area of Science:
- Computational linguistics and biomedical informatics.
- Development and evaluation of natural language processing (NLP) algorithms for clinical text analysis.
Background:
- Manual curation of phenotypic concepts (e.g., Human Phenotype Ontology terms) from electronic health records (EHRs) is laborious and prone to errors.
- Natural language processing (NLP) offers automated solutions for phenotype extraction, improving efficiency in clinical data analysis.
- Existing NLP systems may lack portability across different patient cohorts; ensemble methods are explored to enhance this.
Purpose of the Study:
- To compare the performance of individual NLP systems and ensemble techniques for extracting generic and patient-specific phenotypic concepts.
- To evaluate the portability and reproducibility of NLP pipelines across different clinical datasets.
- To introduce a novel evaluation metric accounting for concept granularity, hierarchies, and frequencies.
Main Methods:
- Four NLP systems (MetaMapLite, MedLEE, ClinPhen, cTAKES) and four ensemble techniques (intersection, union, majority-voting, machine learning) were evaluated.
- Performance was assessed on gold-standard datasets annotated by clinical experts, including EHR notes and PubMed case report abstracts.
- A novel evaluation metric was developed to address concept granularity differences.
Main Results:
- Ensemble methods, particularly union and majority-voting, generally outperformed individual NLP systems in generic phenotypic concept recognition across datasets.
- Majority-vote ensemble achieved top performance in patient-specific phenotype identification in both tested EHR datasets.
- Individual NLP systems showed varied performance, often excelling on datasets they were primarily designed for, highlighting the benefit of ensembles for broader applicability.
Conclusions:
- Ensembles of NLP systems significantly enhance both generic and patient-specific phenotypic concept identification compared to standalone systems.
- Ensemble approaches improve the reproducibility of NLP results across different cohorts and tasks, offering a more portable phenotyping solution.
- The findings support the use of ensemble NLP methods for more robust and generalizable automated clinical phenotyping.
Related Concept Videos
Background and Environment Affect Phenotype
An example of how genetic background affects phenotype can be seen in horses. The Extension gene in horses is responsible for their coat color. A wild-type gene (EE) produces black pigment in the coat, while a mutant gene (ee) produces red pigment. A...
Automatic Processing and Automatic Social Behavior

