Related Experiment Video
Updated: Jun 25, 2026

Behavioral Assessment of Visual Function via Optomotor Response and Cognitive Function via Y-Maze in Diabetic Rats
Published on: October 23, 2020
Analysis and Comparison of Two Artificial Intelligence Diabetic Retinopathy Screening Algorithms in a Pilot Study:
Andrzej Grzybowski1,2, Piotr Brona3
1Department of Ophthalmology, University of Warmia and Mazury, Żołnierska 18, 10-561 Olsztyn, Poland.
This study compared two automated artificial intelligence systems for detecting diabetic retinopathy. Researchers evaluated how well these tools matched an ophthalmologist's assessment using images from a local screening program. The findings highlight differences in performance between the systems and suggest that further research is needed to guide clinical selection.
Area of Science:
- Ophthalmology research regarding diabetic retinopathy screening algorithms
- Artificial intelligence in medical diagnostics
Background:
The rising incidence of diabetic retinopathy creates significant pressure on existing medical infrastructure. Automated diagnostic tools offer a potential solution to manage this growing patient volume. No prior work had resolved the performance differences between various commercially available screening platforms. That uncertainty drove the need for comparative evaluations of these emerging technologies. Prior research has shown that autonomous systems can identify retinal disease patterns effectively. However, limited data exist regarding how these diverse software solutions perform against one another. This gap motivated a head-to-head assessment of two prominent diagnostic programs. Clinicians currently lack clear guidance when selecting the most appropriate software for their specific screening environments.
Purpose Of The Study:
The aim of this investigation was to perform a direct comparison between two autonomous screening programs for retinal disease. Researchers sought to address the difficulty of evaluating different commercial solutions currently available to clinicians. The study specifically examined the diagnostic performance of IDx-DR and Retinalyze using a shared set of patient images. This work was motivated by the increasing burden on healthcare resources caused by rising disease prevalence. The team intended to provide clarity for practitioners choosing between various emerging software options. By analyzing a local screening dataset, the authors explored how these tools align with expert manual grading. The project also aimed to identify potential limitations in current evaluation methods for automated diagnostic systems. This effort serves as a pilot to guide more comprehensive future research in the field.
Main Methods:
Review approach involved a retrospective analysis of images collected from a local screening program. The team selected 60 subjects with referable disease and 110 individuals without evidence of pathology. Investigators required four clear, 45-degree field of view captures per participant. Each patient provided both macula-centered and disc-centered images for both eyes. A Topcon NW-400 device generated all visual data without pharmacological pupil dilation. The study compared two specific software solutions against a single expert's manual classification. Researchers tested two distinct operational configurations for one of the programs to evaluate sensitivity. This methodology allowed for a direct assessment of diagnostic consistency across different algorithmic strategies.
Main Results:
Key findings from the literature indicate that IDx-DR achieved 93.3% agreement for positive cases and 95.5% for negative cases. Retinalyze strategy 1 demonstrated 89.7% agreement for positive subjects and 71.8% for negative subjects. Under strategy 2, Retinalyze showed 74.1% agreement for positive cases and 93.6% for negative cases. Both diagnostic platforms successfully processed the vast majority of the provided image sets. The investigators observed that both systems were straightforward to configure and operate within the clinical workflow. Performance metrics varied notably depending on the specific configuration strategy applied to the software. These results suggest that system settings significantly impact the diagnostic output for retinal disease detection. The data provide an initial baseline for comparing these two commercial screening technologies.
Conclusions:
The authors suggest that performance variations exist between the two tested diagnostic platforms. These findings indicate that software configuration significantly influences the sensitivity and specificity of automated screening results. The researchers propose that future investigations should utilize larger, more representative patient cohorts to validate these initial observations. Synthesis and implications reveal that current reference grading methods require refinement to ensure high-quality comparisons. The study highlights that both systems remain user-friendly despite observed differences in diagnostic accuracy. The authors emphasize that clinicians must carefully consider specific operational strategies when deploying these tools. The team notes that the pilot nature of this work limits the generalizability of the reported metrics. Future efforts should prioritize standardized protocols to improve the robustness of automated retinal disease detection.
Frequently Asked Questions
The researchers found that IDx-DR achieved 93.3% agreement for positive cases and 95.5% for negative cases. In contrast, Retinalyze strategy 1 reached 89.7% and 71.8%, while strategy 2 attained 74.1% and 93.6% agreement respectively.
The team utilized a Topcon NW-400 fundus camera to capture four images per subject. This included both macula-centered and disc-centered views for each eye, all obtained without the use of mydriatic drops to dilate the pupils.
The authors state that high-quality images are necessary because the software requires four distinct fields of view to function. If images lack sufficient clarity or coverage, the automated programs cannot reliably assess the retinal status of the patient.
The researchers employed a retrospective design using a non-representative sample of 60 positive and 110 negative subjects. This data set served as the basis for comparing the automated outputs against a single ophthalmologist's manual grading.
The investigators measured the percentage agreement between the software outputs and a single ophthalmologist's reference grade. They also assessed the ease of system setup and the ability of the programs to successfully analyze the provided image sets.
The researchers propose that the limitations regarding sample selection and reference grading must be addressed. They suggest that these improvements are required before conducting a more robust, large-scale investigation into these diagnostic tools.

