Related Experiment Video
Updated: Nov 22, 2025

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Multicenter, Head-to-Head, Real-World Validation Study of Seven Automated Artificial Intelligence Diabetic
Aaron Y Lee1,2,3, Ryan T Yanagihara4, Cecilia S Lee4,2
1Department of Ophthalmology, University of Washington School of Medicine, Seattle, WA leeay@uw.edu.
This study compared seven automated artificial intelligence systems for detecting diabetic retinopathy using over 300,000 real-world retinal images. While some systems showed promise, performance varied significantly, highlighting the need for rigorous testing before these tools are used in primary care.
Area of Science:
- Ophthalmology diagnostics and Diabetic Retinopathy screening research
- Artificial intelligence in clinical medicine
Background:
No prior work had resolved the comparative effectiveness of various automated diagnostic tools for eye disease in routine clinical environments. That uncertainty drove the need for systematic evaluation of these emerging technologies. Prior research has shown that rising global disease prevalence necessitates efficient screening solutions. While some systems possess regulatory clearance, many others operate without comprehensive validation against human experts. This gap motivated a multicenter assessment of multiple algorithms using large-scale patient datasets. Existing literature often lacks head-to-head comparisons of these diverse digital platforms. Most previous investigations focused on isolated performance metrics rather than broad, real-world application. Consequently, the clinical community remains divided on the readiness of these automated systems for widespread deployment.
Purpose Of The Study:
The primary aim was to compare the performance of seven automated screening algorithms against human graders using real-world imaging data. This study addressed the lack of systematic validation for various diagnostic tools currently in clinical use. Researchers sought to determine if these systems could reliably identify referable disease in primary care environments. The motivation stemmed from the increasing global burden of eye complications associated with diabetes. While some platforms have received regulatory clearance, their actual accuracy in diverse clinical settings remained unverified. The team intended to provide clear evidence regarding the readiness of these technologies for widespread adoption. By analyzing a large-scale veteran population, the authors aimed to establish a benchmark for future diagnostic development. This work clarifies the current landscape of automated screening to inform better clinical decision-making.
Main Methods:
The investigation employed a multicenter, noninterventional design to assess diagnostic performance. Researchers analyzed a vast collection of retinal scans obtained from two major health care facilities. Five distinct commercial entities provided the seven algorithms tested in this head-to-head comparison. All systems processed the entire image set independently, regardless of initial scan quality. The team compared automated outputs against both original human grades and a secondary arbitrated dataset. This approach ensured a comprehensive evaluation of sensitivity and specificity for referable disease detection. Furthermore, the investigators calculated the estimated value per encounter to assess economic implications. The study period spanned twelve years, capturing a wide range of clinical scenarios and patient demographics.
Main Results:
The strongest finding indicates that sensitivities among the seven algorithms varied widely, ranging from 50.98% to 85.90%. Although high negative predictive values of 82.72% to 93.69% were observed, most systems did not surpass human performance. One specific algorithm achieved comparable sensitivity of 80.47% and specificity of 81.28% against the arbitrated dataset. Conversely, one system demonstrated significantly lower sensitivity of 74.42% for proliferative disease compared to human graders. The statistical significance for this performance deficit was recorded at P = 9.77 × 10^-4. Economic analysis showed that the value per encounter for ophthalmologists ranged between $15.14 and $18.06. For optometrists, the estimated value per encounter was lower, falling between $7.74 and $9.24. These results demonstrate substantial differences in both diagnostic accuracy and potential cost efficiency across the platforms.
Conclusions:
The authors propose that significant performance variability exists among the tested automated diagnostic systems. These findings suggest that current algorithms do not uniformly outperform human experts in clinical settings. The researchers emphasize that rigorous testing on diverse, real-world datasets is necessary before widespread implementation. One specific system demonstrated comparable sensitivity and specificity to human graders in this analysis. However, another tool showed inferior detection capabilities for proliferative disease compared to human clinicians. The study highlights that negative predictive values remain generally high across all evaluated platforms. These results imply that regulatory approval alone does not guarantee superior diagnostic accuracy in practice. The authors conclude that ongoing validation remains critical to ensure patient safety and diagnostic reliability.
Frequently Asked Questions
The researchers report that sensitivities ranged from 50.98% to 85.90% across the seven systems. While one algorithm achieved comparable performance to human graders, others failed to match the diagnostic accuracy of traditional teleretinal screening methods.
The study utilized a massive dataset comprising 311,604 retinal images collected from 23,724 veterans. These scans were obtained from two distinct health care systems between 2006 and 2018 to ensure a robust, multicenter evaluation.
The authors state that comparing these systems against both original teleretinal grades and a regraded arbitrated dataset was necessary. This dual-comparison approach allowed for a more precise determination of how these tools perform relative to human clinical standards.
The researchers used these images to classify cases as either referable disease or non-referable. This binary classification serves as the foundation for determining whether a patient requires further specialist intervention or routine follow-up.
The study measured the value per encounter, which ranged from $15.14 to $18.06 for ophthalmologists and $7.74 to $9.24 for optometrists. This economic metric provides insight into the potential cost efficiency of integrating these tools into existing workflows.
The authors argue that their results support the necessity of rigorous, real-world testing for all such technologies. They suggest that clinical implementation should only proceed after these systems demonstrate consistent reliability compared to established human diagnostic benchmarks.

