Related Experiment Video
Updated: Aug 11, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Leveraging Natural Language Processing to Extract Features of Colorectal Polyps From Pathology Reports for
Ryzen Benson1, Candace Winterton2, Maci Winn2,3
1Department of Biomedical Informatics, University of Utah, Salt Lake City, UT.
A new natural language processing (NLP) pipeline accurately extracts colorectal polyp histopathology from reports. This enables linking data with electronic health records (EHR) for improved risk factor research.
Area of Science:
- Medical Informatics
- Gastroenterology
- Computational Pathology
Background:
- Histopathologic features of colorectal polyps are crucial for risk factor research.
- Manual data extraction from unstructured pathology reports is labor-intensive and costly.
- Electronic Health Record (EHR) data offers valuable insights but requires integration with detailed pathology findings.
Purpose of the Study:
- To develop and evaluate a natural language processing (NLP) pipeline for automated extraction of colorectal polyp histopathologic features from pathology reports.
- To specifically focus on accurately extracting individual polyp size.
- To create an analysis-ready dataset by linking NLP-extracted features with structured EHR data.
Main Methods:
- Utilized 24,584 colonoscopy pathology reports from the University of Utah.
- Developed an annotation scheme and reference standard using 350 manually annotated reports.
- Evaluated pipeline performance against the reference standard for features including location, histology, size, shape, and dysplasia.
- Applied the validated pipeline to 24,225 unseen reports and integrated the extracted data with EHR data.
Main Results:
- The NLP pipeline achieved high performance with 98.4% F1-score, 98.9% precision, and 98.0% recall.
- Accurate extraction rates for key features included 95.6% for size, 97.2% for location, and 97.8% for histology.
- The pipeline processed 28,387 polyps, identifying tubular adenomas (55.9%) and advanced adenomas (8.1%), with a mean polyp size of 0.57 cm.
Conclusions:
- The developed NLP pipeline accurately extracts critical histopathologic features, including polyp size, from colonoscopy reports.
- This automated approach significantly enhances the efficiency of creating research-ready datasets.
- The integration of NLP-extracted pathology data with EHR data facilitates robust epidemiologic studies on colorectal polyp risk factors and outcomes.
More Related Videos
07:35Evaluation of Colorectal Cancer Risk and Prevalence by Stool DNA Integrity Detection
Published on: June 8, 2020
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025