Related Experiment Video
Updated: Jul 18, 2025

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Comparing Machine Learning Models and Human Raters When Ranking Medical Student Performance Evaluations
Jonathan Kibble1, Jeffrey Plochocki2
1Both authors are with University of Central Florida College of Medicine. is Professor of Medical Education; and.
Background:
The Medical Student Performance Evaluation (MSPE), a narrative summary of each student's academic and professional performance in US medical school is long, making it challenging for residency programs evaluating large numbers of applicants.
Objective:
To create a rubric to assess MSPE narratives and to compare the ability of 3 commercially available machine learning models (MLMs) to rank MSPEs in order of positivity.
Methods:
Thirty out of a possible 120 MSPEs from the University of Central Florida class of 2020 were de-identified and subjected to manual scoring and ranking by a pair of faculty members using a new rubric based on the Accreditation Council for Graduate Medical Education competencies, and to global sentiment analysis by the MLMs. Correlation analysis was used to assess reliability and agreement between student rank orders produced by faculty and MLMs.
Results:
The intraclass correlation coefficient used to assess faculty interrater reliability was 0.864 (P<.001; 95% CI 0.715-0.935) for total rubric scores and ranged from 0.402 to 0.768 for isolated subscales; faculty rank orders were also highly correlated (rs=0.758; P<.001; 95% CI 0.539-0.881). The authors report good feasibility as the rubric was easy to use and added minimal time to reading MSPEs. The MLMs correctly reported a positive sentiment for all 30 MSPE narratives, but their rank orders produced no significant correlations between different MLMs, or when compared with faculty rankings.
Conclusions:
The rubric for manual grading provided reliable overall scoring and ranking of MSPEs. The MLMs accurately detected positive sentiment in the MSPEs but were unable to provide reliable rank ordering.
More Related Videos
04:54Author Spotlight: IntelliSleepScorer — A High-Accuracy, Accessible GUI Software for Automated Sleep Stage Scoring in Mice and its Application in Psychiatric Research
Published on: November 8, 2024
12:51Evaluation of Commercial-Off-The-Shelf Wrist Wearables to Estimate Stress on Students
Published on: June 16, 2018
Related Concept Videos
Ranks
Comparing the Survival Analysis of Two or More Groups
Regression Toward the Mean