Related Experiment Video
Updated: Dec 13, 2025

09:20
Cloud-Based Phrase Mining and Analysis of User-Defined Phrase-Category Association in Biomedical Publications
Published on: February 23, 2019
9.1K
Experiences implementing scalable, containerized, cloud-based NLP for extracting biobank participant phenotypes at
Timothy A Miller1,2, Paul Avillach1,3, Kenneth D Mandl1,2
1Computational Health Informatics Program, Boston Children's Hospital, Boston, Massachusetts, USA.
JAMIA Open
|August 1, 2020
Summary
Scalable natural language processing (NLP) infrastructure efficiently processes electronic health records (EHRs), extracting millions of concepts for comprehensive patient insights and high-throughput phenotyping.
Area of Science:
- Biomedical Informatics
- Computational Linguistics
Background:
- Electronic health records (EHRs) contain vast amounts of unstructured free text.
- Extracting meaningful information from EHR free text is crucial for clinical research and patient care.
Purpose of the Study:
- To develop scalable natural language processing (NLP) infrastructure for processing EHR free text.
- To enhance the utility of EHR data through efficient concept extraction.
Main Methods:
- Extended open-source Apache cTAKES NLP software with standard scalability technologies.
- Monitored component queue size to identify and resolve processing bottlenecks.
- Processed EHR free text data from the PrecisionLink Biobank.
Main Results:
- Processed over 1.2 million notes for more than 8000 patients.
- Extracted 154 million concepts from the EHR data.
- Achieved processing speeds exceeding 1 million notes per day with optimized configurations.
Conclusions:
- Scalable NLP infrastructure enables efficient processing of large EHR document collections.
- Extracted NLP concepts provide a more complete understanding of patient status.
- This approach supports high-throughput phenotyping for clinical research.

