Related Experiment Video
Updated: Apr 7, 2026

08:27
Three-Dimensional Finger Motion Tracking during Needling: A Solution for the Kinematic Analysis of Acupuncture Manipulation
Published on: October 28, 2021
3.3K
A study of crowdsourced segment-level surgical skill assessment using pairwise rankings
Anand Malpani1, S Swaroop Vedula, Chi Chiung Grace Chen
1Johns Hopkins University, 3400 N Charles St, Hackerman Hall Room 200, Baltimore, MD, USA, anandmalpani@jhu.edu.
Summary
This study developed a reliable framework for objective surgical skill assessment using crowdsourced segment ratings. The method provides valid, segment-specific feedback for surgical training, improving technical skill evaluation.
Area of Science:
- Medical Education Technology
- Surgical Skill Assessment
- Human-Computer Interaction
Background:
- Current surgical skill assessment methods are subjective or offer only global task evaluations.
- Global evaluations lack specific feedback, hindering trainee improvement in task segments.
- Objective, segment-specific assessment is needed for effective surgical training.
Purpose of the Study:
- To investigate the reliability and validity of a framework for objective, segment-level surgical skill assessment.
- To compare assessments from crowdsourced ratings (untrained individuals and experts) with manual global ratings.
- To evaluate the framework's ability to provide actionable feedback for surgical trainees.
Main Methods:
- Developed a framework using a binary classifier for pairwise segment preferences.
- Computed segment-level percentile scores based on classifier preferences.
- Predicted task-level scores from segment scores and conducted a crowdsourcing study with untrained individuals and experts for validation.
Main Results:
- Moderate inter-rater reliability observed in both crowd (κ = 0.41) and expert (κ = 0.55) groups.
- Automated classifiers achieved high accuracy (crowd 85%, expert 89%) compared to inter-rater agreement.
- Framework-predicted task scores highly correlated with ground truth (RMSE < 1 SD) and showed strong agreement between crowd and expert assessments (ρ ≥ 0.84).
Conclusions:
- The crowdsourced pairwise comparison framework provides valid objective assessment for surgical skill segments and overall tasks.
- Crowdsourcing offers an efficient and reliable method for segment-level skill comparison in surgical training.
- The framework is suitable for deployment in surgical training for standardized, automated technical skill evaluation.

