Related Experiment Video
Updated: Mar 29, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Current capabilities of large language models as peer reviewers for manuscripts submitted to ophthalmology-related
Majid Moshirfar1,2,3, Kenneth D Han4, Muhammed A Jaafar4
1Hoopes Moshirfar Research Center, Hoopes Vision, Draper, UT.
Abstract:
To compare large language models (LLMs) and human reviewers in the peer review process of manuscripts submitted to 3 ophthalmology-related journals. This retrospective study comprised 300 randomly selected manuscripts from 3 anonymized journals under 1 editor between June 2023 and July 2024. Comments from 2 LLMs (Chat Generative Pre-Trained Transformer [ChatGPT] 4o and Gemini) and human reviewers (324 ophthalmologists) were compared. LLMs were prompted to accept, accept with major or minor revisions, or reject each manuscript in addition to providing comments. A 5-point Likert scale was used to assess the "favorability" of comments and compare manuscripts that were accepted or rejected by the editor. A 4-category quality assessment was used to compare the number of comments, detail/specificity, critical analysis, and literature support. Human reviewers rejected manuscripts more frequently (73.33% vs 2.00% ChatGPT and 2.00% Gemini; P < .001) and suggested major (22.67% vs 68.00% ChatGPT and 31.33% Gemini; P < .001) or minor revisions (3.33% vs 30.00% ChatGPT and 66.33% Gemini; P < .001) less often. Human reviewers gave more negative feedback for rejected manuscripts (-1.05 vs -0.02 ChatGPT and 0.24 Gemini; P < .015). ChatGPT repeated "novelty," "sample size," and "clarity" in 75%, 60%, and 50% of cases, respectively, while Gemini did so in 80%, 70%, and 65% of cases. Both lacked specificity, omitting line numbers and references. Although it is hoped that LLMs will one day be able to augment the role of peer reviewers, in their current state, LLMs should not be used for manuscript revision.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:08Using Optical Coherence Tomography and Optokinetic Response As Structural and Functional Visual System Readouts in Mice and Rats
Published on: January 10, 2019
Related Concept Videos
The Retina
Vision
Proofreading
Genetic Lingo
Glaucoma: Overview
Feedback Inhibition