Related Experiment Video
Updated: May 5, 2026

In vivo Structural Assessments of Ocular Disease in Rodent Models using Optical Coherence Tomography
Published on: July 24, 2020
Large language model-based scribing tools for ophthalmology: performance and safety evaluation using simulated
Minjie Zou1,2, Sahana Srinivasan1,2, Victoria Chin3
1Centre for Innovation and Precision Eye Health, Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore.
Heidi AI led in performance among six LLM scribing tools for ophthalmology notes, but all tools had errors, necessitating human oversight for patient safety.
Area of Science:
- Ophthalmology
- Medical Informatics
- Artificial Intelligence
Background:
- Six large language model (LLM)-based scribing tools were evaluated for their ability to generate ophthalmological clinical notes.
- The study focused on performance and safety in simulated physician-patient encounters.
Purpose of the Study:
- To compare the performance and safety of six LLM-based scribing tools in generating ophthalmological clinical notes.
- To identify the best-performing tool and assess the types and frequency of errors made by these AI scribes.
Main Methods:
- A cross-sectional comparative study used seven ophthalmic encounters transcribed by six scribing tools.
- Note quality was assessed by consultant ophthalmologists using the Physician Documentation Quality Instrument (PDQI-9) and text-generation metrics (ROUGE-L, BERTScore, BARTScore, AlignScore).
- Safety analysis identified documentation errors, including omissions, incorrect information, and extraneous content.
Main Results:
- Heidi AI achieved the highest overall documentation quality (PDQI-9: 38.524) and led in accuracy, usefulness, organization, comprehensiveness, and succinctness.
- Heidi AI also led in ROUGE-L (0.326), BERTScore (0.727), and AlignScore (0.704).
- All tools exhibited errors, primarily omissions of critical clinical details, but also incorrect or extraneous information.
Conclusions:
- Heidi AI demonstrated the highest performance among the evaluated LLM scribing tools for ophthalmology documentation.
- Despite promising performance, all tools exhibited errors, particularly omissions of critical details, posing potential risks to patient care.
- Human oversight and rigorous verification are essential before clinical integration of these AI scribing tools.
More Related Videos
05:49Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
Published on: November 1, 2024
07:12Development of a Gaze-Contingent Display Framework Designed for Perceptual and Oculomotor Research with Simulated Central Vision Loss
Published on: April 11, 2025