Related Experiment Video
Updated: Jul 10, 2026

07:50
A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
16.7K
Benchmarking GPT-5, LLaMA, and Mistral for Clinical Named Entity Recognition in Ophthalmology Progress Notes.
Iyad Majid1, Franklin Y Ruan1, Sophia Y Wang1
1Stanford University School of Medicine, Palo Alto, CA, USA.
Translational Vision Science & Technology
|April 16, 2026
Summary
GPT-5 demonstrated superior performance in extracting clinical data from ophthalmology notes compared to LLaMA 3.1-70B and Mistral 7B. This advancement shows the potential of large language models in improving ophthalmic patient care.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing in Healthcare
- Ophthalmology Data Science
Background:
- Unstructured clinical notes pose challenges for data extraction.
- Large language models (LLMs) show promise for processing medical text.
- Efficient extraction of ophthalmic information is crucial for patient care and research.
Purpose of the Study:
- To compare the performance of GPT-5, LLaMA 3.1-70B, and Mistral 7B in extracting clinically relevant entities from unstructured ophthalmology progress notes.
- To evaluate the accuracy of different LLMs in identifying specific ophthalmic data points.
Main Methods:
- Four hundred eighty deidentified ophthalmology progress notes were used for evaluation.
- GPT-5, LLaMA 3.1-70B (quantized and full-precision), and Mistral 7B were tested with a standardized prompt.
- Outputs were manually annotated for six entity types, and performance was measured using precision, recall, F1 score, and accuracy.
Main Results:
- GPT-5 achieved superior performance across all metrics, with the highest F1 score for surgical history (0.929).
- LLaMA 3.1-70B and Mistral 7B showed strengths in intraocular pressure (IOP) extraction but struggled with diagnostic tests.
- Micro- and macro-averaged F1 scores confirmed GPT-5's overall superiority (0.941 micro, 0.936 macro).
Conclusions:
- LLMs, particularly GPT-5, show significant potential for accurate clinical information extraction from ophthalmic text.
- These models can streamline information retrieval and enhance clinical decision-making in ophthalmology.
- Structured data generation from LLM analysis can accelerate research and improve patient outcomes.

