Related Experiment Video
Updated: Mar 4, 2026

15:30
A Telemetric, Gravimetric Platform for Real-Time Physiological Phenotyping of Plant–Environment Interactions
Published on: August 5, 2020
12.7K
HTP-NLP: A New NLP System for High Throughput Phenotyping
Daniel R Schlegel1, Chris Crowner2, Frank Lehoullier2
1Department of Computer Science, SUNY Oswego, Oswego, NY, USA.
Studies in Health Technology and Informatics
|April 21, 2017
Summary
This study introduces two advances in the High Throughput Phenotyping Natural Language Processing (NLP) system to speed up clinical data processing for research cohort extraction. These improvements enable faster, more efficient analysis of large clinical datasets.
Area of Science:
- Natural Language Processing (NLP)
- Clinical Informatics
- Biomedical Data Science
Background:
- Secondary use of clinical data is crucial for research.
- Efficient cohort extraction from clinical data is a bottleneck.
- Current methods lack the speed for high-throughput analysis.
Purpose of the Study:
- To present advances in the High Throughput Phenotyping NLP system.
- To enable truly high-throughput processing of clinical data.
- To support rapid cohort extraction for researchers.
Main Methods:
- Developed semantic indexing for storing and generalizing partially-processed results.
- Implemented compositional expressions to handle ungrammatical clinical text.
- Evaluated system performance with initial timing results.
Main Results:
- The High Throughput Phenotyping NLP system demonstrates significant speed improvements.
- Semantic indexing enhances the efficiency of data processing.
- Handling of ungrammatical text improves data extraction accuracy.
Conclusions:
- The presented advances significantly enhance the throughput of clinical data processing.
- The improved NLP system facilitates faster and more accurate cohort extraction.
- These developments pave the way for large-scale clinical research using secondary data.

