Related Experiment Video
Updated: Aug 25, 2026

A Training and Testing System for Performing Vascular Reconstruction In Vitro
Published on: October 26, 2019
Reproducibility of Tear Ferning Test Classification by Human Examiners and Artificial Intelligence Models: A
Sérgio Felberg1,2, Guilherme D'Agosto Bernardes1, Felipe Augusto Casseb Dos Santos1
1Department of Ophthalmology, Santa Casa de Misericórdia de São Paulo, São Paulo, SP, Brazil.
Purpose:
To compare the reproducibility of tear ferning classification by human evaluators and 3 multimodal artificial intelligence platforms using the 4-grade Rolando scale and a predefined binary scheme, and to clarify its clinical role.
Methods:
Eighty polarized-light microscopy tear-film images were independently graded (Rolando grades I-IV) by 3 ocular surface researchers and 3 platforms (ChatGPT, Claude, Gemini), each evaluated with an identical prompt and image order in single-run conditions. Grades I-II were defined a priori as normal and III-IV as abnormal.
Results:
For the 4-grade scale, Fleiss kappa was 0.65 (humans), 0.44 (6 raters), and 0.37 (platforms); binary reclassification raised these to 0.76, 0.66, and 0.66. Against consensus, Claude agreed best (binary agreement 93.8%, kappa 0.87; 4-grade 82.5%, weighted kappa 0.84, 95% confidence interval, 0.75-0.91), within the interhuman range. ChatGPT and Gemini showed lower 4-grade agreement (weighted kappa 0.48 each) and significantly underclassified abnormalities (McNemar P = 0.0004 and 0.00006), whereas Claude showed no such asymmetry (P = 0.3750).
Conclusions:
Binary classification was more reproducible than the 4-grade Rolando scale for humans and platforms, but it is a pragmatic standardization strategy, not a replacement for ordinal grading. Claude best matched the human benchmark, whereas ChatGPT and Gemini were less reproducible and systematically underclassified abnormalities. Platform assistance may have adjunctive value, but clinical use will require external validation, accuracy and repeated-session testing, and safety evaluation.
