Related Experiment Video
Updated: Jul 13, 2026

Author Spotlight: 3D Scanning and Augmented Reality for Enhanced Cancer Surgery Communication
Published on: December 15, 2023
The utility of including pathology reports in improving the computational identification of patients
Wei Chen1, Yungui Huang1, Brendan Boyle2
1Department of Research and Development, Research Information Solutions and Innovation, Nationwide Children's Hospital, 575 Children's Crossroad, Columbus, Ohio 43215, USA.
Insights
Automated classification using pathology reports and clinical data significantly improves celiac disease (CD) patient identification compared to ICD-9 codes alone. This approach enhances diagnostic accuracy for better chronic disease management.
Area of Science:
- Medical Informatics
- Immunology
- Gastroenterology
Background:
- Celiac disease (CD) is a prevalent autoimmune disorder requiring accurate patient identification for effective management.
- Current methods relying solely on International Classification of Diseases-9 (ICD-9) codes are insufficient for precise CD case detection.
- Electronic Health Records (EHRs) offer a rich data source for developing improved identification strategies.
Purpose of the Study:
- To develop and evaluate automated classification algorithms for refining celiac disease patient identification.
- To leverage pathology reports and clinical data within EHRs to enhance accuracy beyond traditional ICD-9 coding.
- To compare the performance of machine learning models against ICD-9 code-based identification.
Main Methods:
- Utilized EHR data, including ICD-9 codes (579.0) and tissue transglutaminase laboratory results.
- Applied natural language processing (NLP) to analyze pathology reports from upper endoscopies.
- Trained and evaluated twelve machine learning classifiers using a combination of clinical and pathology data, employing ten-fold cross-validation.
Main Results:
- A logistic model incorporating clinical and pathology report features achieved superior performance (Kappa: 0.78, F1: 0.92, AUC: 0.94).
- In contrast, using ICD-9 codes alone yielded significantly lower performance metrics (Kappa: 0.28, F1: 0.75, AUC: 0.63).
- The study analyzed 1498 patient records, comprising 363 confirmed celiac disease cases and 1135 false positives.
Conclusions:
- The developed automated classification system offers an efficient and reliable method for improving celiac disease patient identification.
- Integrating pathology report analysis with clinical data enhances diagnostic accuracy in EHRs.
- This advanced approach holds promise for optimizing the management of celiac disease.
Background:
Celiac disease (CD) is a common autoimmune disorder. Efficient identification of patients may improve chronic management of the disease. Prior studies have shown searching International Classification of Diseases-9 (ICD-9) codes alone is inaccurate for identifying patients with CD. In this study, we developed automated classification algorithms leveraging pathology reports and other clinical data in Electronic Health Records (EHRs) to refine the subset population preselected using ICD-9 code (579.0).
Materials And Methods:
EHRs were searched for established ICD-9 code (579.0) suggesting CD, based on which an initial identification of cases was obtained. In addition, laboratory results for tissue transglutaminse were extracted. Using natural language processing we analyzed pathology reports from upper endoscopy. Twelve machine learning classifiers using different combinations of variables related to ICD-9 CD status, laboratory result status, and pathology reports were experimented to find the best possible CD classifier. Ten-fold cross-validation was used to assess the results.
Results:
A total of 1498 patient records were used including 363 confirmed cases and 1135 false positive cases that served as controls. Logistic model based on both clinical and pathology report features produced the best results: Kappa of 0.78, F1 of 0.92, and area under the curve (AUC) of 0.94, whereas in contrast using ICD-9 only generated poor results: Kappa of 0.28, F1 of 0.75, and AUC of 0.63.
Conclusion:
Our automated classification system presented an efficient and reliable way to improve the performance of CD patient identification.
More Related Videos
07:32Author Spotlight: Investigating Immune Cell Dynamics in the Tumor Microenvironment — Challenges and Innovations in Cancer Prognosis
Published on: April 12, 2024
05:33Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025