Related Experiment Video
Updated: May 2, 2026

09:17
Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
22.2K
Amending AI Software Accuracy for Diabetic Retinopathy Detection Using Conditional Probability and the Appropriate
Christian Segovia1, Daniela Salinas-Toro2, Carla Moraga2
1Centro de Investigaciones y Estudios Avanzados del Maule, Universidad Católica del Maule, Talca, Chile.
Summary
The automated diabetic retinopathy screening tool DART (Diagnóstico Automatizado de Retinografías Telemáticas) showed lower sensitivity and specificity than initially reported. Expert ophthalmologist retinography outperformed DART in detecting diabetic retinopathy.
Area of Science:
- Ophthalmology
- Artificial Intelligence in Medicine
- Medical Diagnostics
Background:
- Diabetic retinopathy (DR) screening is crucial for early detection and prevention of vision loss.
- Automated AI-based tools like DART (Diagnóstico Automatizado de Retinografías Telemáticas) are being implemented in healthcare systems.
- Accurate validation of AI diagnostic tools is essential for reliable clinical decision-making.
Purpose of the Study:
- To re-evaluate and correct the reported sensitivity and specificity of the DART AI tool for diabetic retinopathy detection.
- To compare the performance of DART against traditional methods using appropriate gold standards and conditional probability.
- To assess the reliability of DART within the Chilean public healthcare system.
Main Methods:
- Conditional probability was applied to clinical validation data of DART.
- Three hypothetical sensitivity levels (90%, 80%, 70%) for expert retinography were used to calculate corrected DART performance metrics.
- Performance metrics included sensitivity, specificity, false negative/positive rates (%FN/%FP), and predictive values (NPV/PPV).
Main Results:
- Corrected sensitivity and specificity values for DART were significantly lower than initially reported.
- Expert ophthalmologist retinography (method 2) consistently outperformed AI-based DART (method 3) across all evaluated metrics.
- This includes superior performance in false negative/positive rates and predictive values for expert retinography.
Conclusions:
- The study highlights the critical need for rigorous validation of AI-based diagnostic tools using appropriate gold standards.
- Initial performance reports for DART may have overestimated its diagnostic accuracy for diabetic retinopathy.
- Trustworthy diagnostic parameters are fundamental for the safe and effective integration of AI in clinical practice.

