You might also read
Articles linked to this work by shared authors, journal, and citation graph.
Updated: Jun 25, 2026

Using Retinal Imaging to Study Dementia
Published on: November 6, 2017
Luis Filipe Nakayama1, Lucas Zago Ribeiro2, Mariana Batista Gonçalves2,3,4
1Physician, Department of Ophthalmology, Universidade Federal de São Paulo - EPM, Botucatu Street, 821, Vila Clementino, São Paulo, SP, 04023-062, Brazil. nakayama.luis@gmail.com.
This review examines how different grading systems for diabetic retinopathy affect the performance of machine learning models. It highlights the challenges caused by inconsistent labeling standards in medical image datasets and suggests ways to improve data reliability for better diagnostic tools.
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
04:36Author Spotlight: Understanding Retinal Vessel Resilience and Disease Progression
Published on: January 12, 2024
Area of Science:
Background:
No prior work has fully resolved the challenges posed by inconsistent grading standards in retinal imaging datasets. Prior research has shown that automated diagnostic tools offer significant improvements in screening efficiency and clinical workflow. That uncertainty drove the need to evaluate how diverse classification methods impact model performance. It was already known that high-quality annotated data remains a primary requirement for training robust supervised learning systems. This gap motivated a closer look at the current landscape of labeling protocols used globally. Researchers have long struggled to reconcile different clinical guidelines when building large-scale medical databases. The absence of a unified ground truth standard complicates the development of reliable automated screening solutions. This review addresses the critical need for consistency in how retinal findings are categorized for computational analysis.
Purpose Of The Study:
This review aims to evaluate the impact of various grading systems on the performance of supervised machine learning algorithms for retinal disease. The authors seek to address the challenges caused by the absence of standardized ground truth in medical imaging datasets. This study explores how different classification frameworks influence the reliability of automated diagnostic tools. The researchers investigate the limitations inherent in current manual annotation practices used for training artificial intelligence. By comparing established grading scales, the paper identifies the primary obstacles to achieving high-quality data. The motivation for this work stems from the need to improve the accuracy and consistency of computer-aided screening. This review provides a critical assessment of how labeling variability affects the development of robust diagnostic models. The authors intend to propose potential solutions to enhance the trustworthiness of datasets used in ophthalmology.
Main Methods:
The review approach involved a systematic comparison of established grading systems used in retinal imaging research. Investigators examined publicly available datasets to identify common labeling practices and their associated limitations. This analysis focused on the Early Treatment Diabetic Retinopathy Study and the International Clinical Diabetic Retinopathy scales. The review approach also evaluated the National Health Service and the Scottish Diabetic Retinopathy Grading Scheme. Researchers assessed how these diverse frameworks influence the quality of ground truth data for supervised algorithms. The study utilized a comparative framework to highlight discrepancies between manual annotation methods. This approach allowed for the identification of systemic issues in current data preparation workflows. The investigators synthesized findings to propose potential pathways for achieving greater consistency in future medical image classification.
Main Results:
Key findings from the literature reveal that the lack of unified ground truth standards represents a primary obstacle for supervised artificial intelligence. The analysis shows that diverse labeling systems, such as the ETDRS and ICDR, generate significant variability in dataset annotations. The review indicates that this heterogeneity directly impacts the diagnostic accuracy and risk stratification capabilities of automated models. Findings suggest that current manual annotation methods often fail to provide the consistency required for high-quality machine learning training. The literature demonstrates that inconsistent grading protocols hinder the effective comparison of different diagnostic algorithms across studies. Results highlight that direct retinal-finding identification may offer a more reliable alternative to traditional broad-scale grading. The synthesis shows that datasets with more trustworthy labeling are essential for improving the performance of supervised learning systems. The authors report that standardization remains the most effective strategy for overcoming these persistent data quality challenges.
Conclusions:
The authors suggest that standardizing classification protocols is a primary requirement for improving machine learning performance in ophthalmology. They propose that direct identification of specific retinal findings may offer a more reliable alternative to traditional grading systems. The review highlights that inconsistent labeling creates significant barriers to the widespread adoption of automated diagnostic tools. Synthesis and implications indicate that future datasets must prioritize high-quality, trustworthy annotations to ensure clinical utility. The researchers emphasize that current grading variations hinder the comparability of different artificial intelligence models. They conclude that adopting uniform standards will facilitate better risk stratification and patient monitoring. The authors suggest that moving toward more granular data labeling could enhance the accuracy of automated screening systems. This synthesis underscores the necessity of addressing data quality to advance the field of computer-aided diagnosis.
The researchers propose that inconsistent grading systems create a fundamental bottleneck for training supervised learning models. By comparing ETDRS, NHS, ICDR, and SDGS standards, they demonstrate how varying criteria lead to disparate ground truth labels, which ultimately limits the diagnostic accuracy of automated screening tools.
The authors describe several established protocols, including the Early Treatment Diabetic Retinopathy Study (ETDRS) and the International Clinical Diabetic Retinopathy (ICDR) scale. These frameworks serve as the primary methods for manual annotation within publicly available retinal imaging datasets.
Standardization is necessary because current datasets lack a unified ground truth. Without consistent labeling, machine learning models struggle to learn accurate patterns, which prevents the reliable identification of disease severity across different clinical environments.
Manual annotation plays a primary role in establishing the ground truth for supervised learning. The authors note that the quality of these human-derived labels directly dictates the effectiveness of the resulting artificial intelligence algorithms.
The researchers identify a phenomenon where diverse labeling systems generate conflicting data points for the same retinal images. This variability complicates the training process and reduces the overall reliability of automated diagnostic predictions.
The authors propose that direct identification of retinal findings, rather than relying solely on broad grading scales, may provide a more robust path forward. They suggest this approach could lead to more trustworthy datasets and improved clinical outcomes.