Related Experiment Videos
Evaluating Artificial Intelligence as a First-Pass Reader in Fundus Photograph Screening: A Multireader Workflow
Jihyeon Baek1, Richul Oh2, Doohyun Park1
1VUNO Inc., Seoul 06541, Republic of Korea.
Abstract:
Background/Objectives: Double reading with arbitration improves diagnostic reliability in fundus screening but requires repeated human interpretation. Whether artificial intelligence (AI) can serve as a first-pass decision source within this workflow, and whether this applies consistently across retinal diseases with differing inter-reader agreement, remains unclear. Methods: In this retrospective study, 6904 color fundus photographs from 2593 patients at a tertiary screening center were analyzed for age-related macular degeneration (AMD), diabetic retinopathy (DR), and retinal vein occlusion (RVO). Three retina specialists independently labeled each image, and an AI system provided binary classifications at a prespecified operating threshold targeting 0.99-sensitivity. In AI-human double reading, the AI and one reader independently interpreted each image, and a second reader arbitrated discordant cases; this was compared with conventional human-human double reading. Three reader combinations were evaluated per disease. Results: AI-human reading required 1.02-1.11 human reads per image versus 2.00-2.10 for human-human reading. For DR and RVO, human-human reading yielded sensitivities of 0.974 and 0.982 and specificities of 0.999 and 1.000, respectively. Across AI-human combinations, sensitivity and specificity did not differ significantly from human-human reading (all p ≥ 0.05; specificity differences ≤0.001). For AMD (human-human sensitivity 0.845, specificity 1.000), AI-human sensitivity varied: two combinations were higher (0.916 and 0.950; both p < 0.001) and one comparable (0.842; p = 0.742). AMD specificity remained ≥0.978. Conclusions: AI-human reading halved human reading volume without significant loss for DR and RVO; AMD varied by configuration. AI use within double-reading workflows should account for disease-specific inter-reader agreement, reader composition, and operating threshold. This was a single-center, retrospective study, prospective external validation in a multicenter setting is warranted before clinical implementation.