Medical exam question difficulty prediction: An analysis of embedding representations, machine-learning approaches,
Shicong Feng1, Tianpeng Zheng2,3, Hao Hang4
1Graduate School of Education, Peking University, Beijing, China.
Introduction:
Item difficulty prediction is crucial for planning and administrating educational assessments, especially those with high-stakes such as medical licensing examinations. The inconsistent findings across existing studies, however, highlight a critical gap in understanding which modeling components are most influential. This research addresses this gap by systematically investigating several key factors hypothesized to affect prediction performance.
Methods:
This study explored the impact of: (1) model domain specificity, (2) input content granularity (e.g. item stem, correct answer, and distractors), (3) embedding dimensionality, and (4) the choice of the machine learning regressor. By selecting a range of embedding models and a series of Machine Learning models to predict the difficulty of 2815 Multiple-Choice Questions sourced from the National Center for Health Professions Education Development.
Results:
Analyses revealed that XGBoost outperformed other counterparts (Mean RMSE = 0.1779), and the use of a domain-specific MedEmbed-small embedding model consistently improved prediction accuracy (Mean RMSE = 0.1860). Notably, using the item stem and the correct answer as input features achieved the best trade-off between predictive accuracy and model parsimony (RMSE = 0.1756).
Discussion:
These findings offer valuable insights for data-driven measurement practices including Automated Item Calibration, Computerized Adaptive Testing, and Intelligent Tutoring Systems in medical education. Furthermore, this study revealed that the optimal feature set for difficulty prediction is contingent on the item style. Future research should extend this line of inquiry to the difficulty prediction of Multimodal test items. [Box: see text].


