Related Experiment Video
Updated: Apr 8, 2026

Quantification of Vascular Parameters in Whole Mount Retinas of Mice with Non-Proliferative and Proliferative Retinopathies
Published on: March 12, 2022
Dealing with inter-expert variability in retinopathy of prematurity: A machine learning approach
V Bolón-Canedo1, E Ataer-Cansizoglu2, D Erdogmus2
1Department of Computer Science, Universidade da Coruña, A Coruña, Spain.
Insights
Disagreements in diagnosing retinopathy of prematurity (ROP) stem from differing expert feature selection. Machine learning identified key features, improving diagnostic consistency and accuracy for this infant eye disease.
Area of Science:
- Ophthalmology
- Medical Imaging
- Machine Learning
Background:
- Inter-expert variability in clinical decision-making, particularly in diagnosing retinopathy of prematurity (ROP), poses a significant challenge.
- ROP affects premature infants and is a leading cause of childhood blindness, highlighting the need for accurate and consistent diagnosis.
- Discrepancies in the features experts consider are a neglected cause of diagnostic variability.
Purpose of the Study:
- To propose and evaluate a machine learning methodology for understanding inter-expert variability in ROP diagnosis.
- To identify key diagnostic features used by experts and assess their impact on diagnostic agreement.
Main Methods:
- Utilized a dataset of 34 retinal images with diagnoses from 22 independent experts.
- Applied feature selection techniques to identify crucial features for each expert.
- Compared feature sets across experts using similarity measures.
- Developed an automated diagnosis system to evaluate the methodology's effectiveness.
Main Results:
- Identified features consistently selected by feature selection methods across different experts.
- Observed a correlation between high expert agreement and similar feature selections.
- Demonstrated that using selected features either improved or maintained the classification performance of the automated system.
Conclusions:
- The proposed methodology effectively identifies expert-relevant features and quantifies inter-expert agreement/disagreement.
- Findings suggest potential for improved diagnostic accuracy and standardization in ROP diagnosis.
- The framework may be applicable to other clinical problems characterized by inter-expert variability.
Background And Objective:
Understanding the causes of disagreement among experts in clinical decision making has been a challenge for decades. In particular, a high amount of variability exists in diagnosis of retinopathy of prematurity (ROP), which is a disease affecting low birth weight infants and a major cause of childhood blindness. A possible cause of variability, that has been mostly neglected in the literature, is related to discrepancies in the sets of important features considered by different experts. In this paper we propose a methodology which makes use of machine learning techniques to understand the underlying causes of inter-expert variability.
Methods:
The experiments are carried out on a dataset consisting of 34 retinal images, each with diagnoses provided by 22 independent experts. Feature selection techniques are applied to discover the most important features considered by a given expert. Those features selected by each expert are then compared to the features selected by other experts by applying similarity measures. Finally, an automated diagnosis system is built in order to check if this approach can be helpful in solving the problem of understanding high inter-rater variability.
Results:
The experimental results reveal that some features are mostly selected by the feature selection methods regardless the considered expert. Moreover, for pairs of experts with high percentage agreement among them, the feature selection algorithms also select similar features. By using the relevant selected features, the classification performance of the automatic system was improved or maintained.
Conclusions:
The proposed methodology provides a handy framework to identify important features for experts and check whether the selected features reflect the pairwise agreements/disagreements. These findings may lead to improved diagnostic accuracy and standardization among clinicians, and pave the way for the application of this methodology to other problems which present inter-expert variability.
More Related Videos
09:28Oxygen-Induced Retinopathy Model for Ischemic Retinal Diseases in Rodents
Published on: September 16, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018