Related Experiment Video
Updated: Feb 28, 2026

Use of a Video Scoring Anchor for Rapid Serial Assessment of Social Communication in Toddlers
Published on: March 14, 2018
Improving Consensus Scoring of Crowdsourced Data Using the Rasch Model: Development and Refinement of a Diagnostic
Christopher John Brady1, Lucy Iluka Mudie1, Xueyang Wang1
1Dana Center for Preventive Ophthalmology, Wilmer Eye Institute, Johns Hopkins University School of Medicine, Baltimore, MD, United States.
Crowdsourcing retinal images for diabetic retinopathy (DR) screening is effective. Using regression analysis to weight worker classifications significantly improves diagnostic accuracy compared to simple majority vote.
Area of Science:
- Ophthalmology
- Medical Imaging
- Computational Biology
Background:
- Diabetic retinopathy (DR) is a primary cause of vision loss in working-age adults globally.
- Current DR screening methods are effective but underutilized, necessitating innovative approaches for detection.
- Crowdsourcing offers a potential avenue for increasing the accessibility and efficiency of DR screening.
Purpose of the Study:
- To evaluate the diagnostic accuracy of crowdsourced gradings of retinal fundus photographs compared to expert grading.
- To determine if regression methods can enhance the consensus accuracy of crowdsourced classifications beyond simple majority vote.
Main Methods:
- 1200 retinal images from the Messidor dataset were classified by 10 trained Amazon Mechanical Turk (AMT) workers each.
- Majority Vote (MV) consensus was established if half or more workers deemed an image abnormal.
- Rasch analysis was employed to calculate worker ability scores, which were then used as weights in a logistic regression model to improve consensus grading accuracy.
Main Results:
- Majority Vote grading achieved 75.5% accuracy, 75.5% sensitivity, 75.5% specificity, and an AUROC of 0.75.
- A logistic regression model incorporating Rasch-weighted worker scores yielded a significantly higher AUROC of 0.91 (95% CI 0.88-0.93).
- Optimizing for 90% sensitivity resulted in 77.5% overall accuracy with 90.3% sensitivity and 68.5% specificity.
Conclusions:
- Crowdsourced interpretation of retinal images provides a rapid and accurate method for DR screening when compared to gold-standard grading.
- Employing logistic regression with Rasch analysis to weight classifications by worker ability substantially improves the accuracy of aggregated diagnostic grades over simple majority vote.
More Related Videos
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Receiver Operating Characteristic Plot
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...