Related Experiment Video
Updated: Nov 16, 2025

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
Published on: December 9, 2022
Using Data Mining for the Early Identification of Struggling Learners in Physician Assistant Education
Erik W Black1,2,3, Shalon R Buchs1,2,3, Breann Garbas1,2,3
1Erik W. Black, PhD, MPH , is an associate professor of Pediatrics and Education at the University of Florida College of Medicine, Gainesville, Florida.
Purpose:
Despite the importance of early intervention and remediation, the relatively short duration of physician assistant education programs necessitates the importance of early identification of at-risk learners. This study sought to ascertain whether machine learning was more effective than logistic regression in predicting remediation status among students, using the limited set of data available before or immediately following the first semester of study as predictor variables and academic remediation as an outcome variable.
Methods:
The analysis included one institution and student data from 177 graduates between 2017 and 2019. We employed one data mining model, random forest trees, and compared it to a traditional predictive analysis method, logistic regression. Due to the small sample size, we employed leave-one-out cross-validation and bootstrap aggregation.
Results:
Data provided evidence that the random forest algorithm correctly identified individuals who would later experience academic intervention with a 63.3% positive predictive value, whereas logistic regression exhibited a positive predictive value of 16.6%.
Conclusions:
This single-institution study indicates that predictive modeling, employing machine learning, may be a more effective means than traditional statistical methods of identifying and providing assistance to learners who may experience academic challenges.

