Evaluating Clinical NLP Services for Chest Radiograph Report Labeling: A Comparative Study on an Independent
Shruti Hegde1, Mabon Manoj Ninan2, Jonathan R Dillman3
1Cincinnati Children's Hospital Medical Center, Cincinnati, OH, USA. Shruti.Hegde@cchmc.org.
Journal of Imaging Informatics in Medicine
|July 29, 2026
Summary
This study evaluated general-purpose clinical natural language processing (NLP) models for labeling pediatric chest X-ray reports, finding significant variability in performance and highlighting the need for further validation before clinical use.
Area of Science:
- Medical Informatics
- Radiology Informatics
- Natural Language Processing
Background:
- Clinical natural language processing (NLP) tools are vital for research and quality improvement by automating clinical report labeling.
- Independent evaluations of these tools for specific tasks, such as pediatric chest X-ray (CXR) report analysis, are limited.
- This study addresses this gap by comparing the performance of commercial NLP models on pediatric CXR reports.
Purpose of the Study:
- To compare the performance of four general-purpose clinical NLP models (AWS, Azure, Google, Spark NLP) against two specialized CXR report labelers (CheXpert, CheXbert).
- To evaluate entity extraction and assertion detection capabilities for clinically relevant findings in pediatric CXR reports.
- To assess the variability in performance across different NLP models and identify areas for improvement.
Main Methods:
- Analysis of 95,008 pediatric CXR reports from a large academic hospital.
- Extraction of entities and assertion statuses (positive, negative, uncertain) from findings and impression sections using four commercial NLP models.
- Comparison of model outputs with specialized labelers (CheXpert, CheXbert) and manual review by a pediatric radiologist on a subset of 360 exams.
- Quantification of inter-model agreement using Fleiss' Kappa and calculation of precision, recall, and F1-scores against ground truth.
Main Results:
- Significant differences observed in the mean number of extracted entities among NLP models (p < 0.001), with Spark NLP extracting the most.
- Assertion distributions also differed significantly (p < 0.001).
- Specialized labelers CheXbert and CheXpert achieved the highest macro F1-scores (0.565 and 0.561) in manual validation, followed by Google (0.453), Azure (0.421), Spark NLP (0.420), and AWS (0.266).
- Performance varied by label, with highest F1 for pleural effusion, pneumothorax, and pneumonia, and lowest for lung opacity, enlarged cardiomediastinum, and lung lesion.
- Positive assertions were best detected (F1 0.782), while uncertain assertions were the most challenging (0.468).
Conclusions:
- Marked variability exists among general-purpose NLP models in entity extraction and uncertainty handling for pediatric CXR reports.
- Specialized CXR NLP tools demonstrated superior performance compared to general-purpose models in this evaluation.
- Further validation and model refinement are crucial before the clinical deployment of NLP tools for pediatric radiology report analysis.

