Related Experiment Video
Updated: Jul 4, 2025

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Using Natural Language Processing to Identify Different Lens Pathology in Electronic Health Records
Joshua D Stein1, Yunshu Zhou2, Chris A Andrews2
1From the W.K. Kellogg Eye Center, Department of Ophthalmology and Visual Sciences, University of Michigan, Ann Arbor, Michigan, USA (J.D.S., Y.Z., C.A.A., J.B.); Department of Health Management and Policy, University of Michigan School of Public Health, Ann Arbor, Michigan, USA (J.D.S.).
Natural language processing (NLP) offers a more accurate method for identifying lens pathology in ophthalmology research compared to traditional billing codes. This advanced NLP algorithm demonstrates high accuracy in classifying ocular conditions from electronic health records.
Area of Science:
- Ophthalmology
- Medical Informatics
- Big Data Analytics
Background:
- Ophthalmology Big Data studies predominantly use International Classification of Diseases (ICD) billing codes.
- ICD codes can be inaccurate or lack specificity for identifying ocular conditions.
- A more precise method is needed to identify lens pathology in large datasets.
Purpose of the Study:
- To assess the accuracy of natural language processing (NLP) in identifying lens pathology.
- To compare NLP's performance against traditional ICD billing codes for lens pathology detection.
Main Methods:
- Developed an NLP algorithm to analyze free-text lens exam data from electronic health records (EHR).
- Applied the algorithm to 17.5 million lens exam records in the Sight Outcomes Research Collaborative (SOURCE) repository.
- Clinician review of 4314 unique lens-exam entries to validate NLP accuracy against ICD codes.
Main Results:
- The NLP algorithm achieved 95.1% accuracy in identifying all lens pathology.
- Sensitivity for NLP was significantly higher (0.98) compared to ICD codes (0.49).
- High clinician agreement for specific conditions like pseudoexfoliation material (100%) and phimosis (99.7%).
Conclusions:
- NLP algorithms can accurately identify and classify lens abnormalities documented in EHRs.
- This approach enhances the ability of researchers to precisely identify ocular pathology.
- NLP facilitates broader and more accurate research using real-world health data.

