Related Experiment Video
Updated: Jul 1, 2026

09:17
Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
21.6K
Grading of diabetic retinopathy using a pre-segmenting deep learning classification model: Validation of an automated
Dyllan Edson Similié1, Jakob K H Andersen2,3, Sebastian Dinesen1
1Department of Ophthalmology, Odense University Hospital, Odense, Denmark.
Acta Ophthalmologica
|October 19, 2024
Summary
An autonomous deep-learning algorithm for diabetic retinopathy (DR) grading showed a high negative predictive value (NPV) of 92% in a real-world population. This suggests its potential clinical feasibility for autonomously identifying patients without DR.
Area of Science:
- Ophthalmology
- Medical Imaging
- Artificial Intelligence
Background:
- Diabetic retinopathy (DR) is a leading cause of vision loss.
- Accurate and efficient DR grading is crucial for timely intervention.
- Deep learning (DL) algorithms offer potential for automated medical image analysis.
Purpose of the Study:
- To validate the performance of an autonomous deep-learning (DL) algorithm for diabetic retinopathy (DR) grading.
- To compare the DL algorithm's performance against a human grader and a gold standard.
- To assess the clinical feasibility of autonomous DR classification.
Main Methods:
- 500 retinal images were graded by an expert ophthalmologist (gold standard) using the International Clinical Diabetic Retinopathy Disease Severity Scale (DR levels 0-4).
- Agreement was measured using weighted kappa for a human grader (with/without DL assistance) and the autonomous DL algorithm.
- Sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were calculated, with discrepancies analyzed.
Main Results:
- The autonomous DL algorithm achieved a weighted kappa of 0.72, sensitivity of 78%, and specificity of 81% compared to the gold standard.
- Extrapolated to a 23.8% real-world DR prevalence, the PPV was 57% and NPV was 92%.
- Discrepancies included artefact detection, missed microaneurysms, and segmentation/classification inconsistencies.
Conclusions:
- The autonomous DL algorithm's performance was comparable to a human grader in some aspects for high-risk populations.
- The high NPV (92%) in a real-world population suggests clinical feasibility for autonomous identification of non-DR patients.
- Further refinement may improve accuracy for DR classification.

