Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Leveraging Large Language Models with Sequential Prompting to Extract Eye Examination Findings from Free-Text
Franklin Y Ruan1, Justin W Lam1, Houri Esmaeilkhanian1
1Department of Ophthalmology, Byers Eye Institute, Stanford University, Palo Alto, California.
Purpose:
The purpose of this study was to evaluate the ability of general domain large language models (LLMs) to identify slit lamp and fundus examination (FE) findings from free-text clinical notes in electronic health records.
Design:
Retrospective cohort study.
Participants:
Data were obtained from the Stanford Research Repository, which contains progress notes of patients from the Department of Ophthalmology at Stanford University since 2008.
Methods:
We assessed a local LLM pipeline, using a novel sequential named entity recognition (NER) prompting approach, for identifying eye examination findings and their lateralities from free-text clinical notes at a single academic center. We compared several local open-source LLMs, including Mistral Nemo 12B, Llama 3.1 8B, and a quantized Llama 3.1 70B. Results were compared to those from a prior study that evaluated Bidirectional Encoder Representations from Transformers (BERT) models fine-tuned on 31 279 weakly labeled ophthalmology progress notes. Testing was conducted on 2 sets: (1) SmartLink-300 templated ophthalmology notes (>5900 entities) and (2) free-text-200 unstructured notes (>3000 entities) from 3 ophthalmologists.
Main Outcome Measures:
Precision, recall, and F1 score for each component in the SLE and FE.
Results:
Llama 3.1 70B was the top-performing model for identifying eye examination findings from SmartLink notes, achieving a microaveraged precision of 0.95, recall of 0.90, and F1 score of 0.92. Llama 3.1 70B performed consistently well across different components of the eye examination, with a microaveraged F1 score of 0.94 for the FE and 0.91 for SLE. For the free-text notes, where fine-tuning a BERT model was not possible, Llama 3.1 70B achieved microaveraged F1 scores of 0.93 with ophthalmologist A notes, 0.94 with ophthalmologist B notes, and 0.96 with ophthalmologist C notes.
Conclusions:
Sufficiently large LLMs, with our sequential prompting approach, outperformed or performed comparably to BERT models fine-tuned on >30 000 examples on the NER task and demonstrated flexibility by performing robustly on data sets BERT could not be fine-tuned on. Our findings suggest LLMs are capable of efficient, automated, and accurate extraction of relevant information from ophthalmology free-text progress notes.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.
More Related Videos
07:36Eye Tracking During Visually Situated Language Comprehension: Flexibility and Limitations in Uncovering Visual Context Effects
Published on: November 30, 2018
09:27Using Eye Movements Recorded in the Visual World Paradigm to Explore the Online Processing of Spoken Language
Published on: October 13, 2018