Related Experiment Video
Updated: Aug 21, 2026

Experimental Autoimmune Uveitis: An Intraocular Inflammatory Mouse Model
Published on: January 12, 2022
Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and
Carlos Cifuentes-González1,2,3, Maxwell B Singer4, William Rojas-Carabali1,2,5
1National Healthcare Group Eye Institute, Tan Tock Seng Hospital, Singapore, Singapore.
Purpose:
To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype.
Design:
Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232).
Subjects:
A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis.
Methods:
Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication.
Main Outcome Measures:
Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement.
Results:
UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous "Need Consensus" outputs (3.4%) compared to experts (16.7%).
Conclusions:
UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology.
Financial Disclosures:
Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.