Medical exam question difficulty prediction: An analysis of embedding representations, machine-learning approaches,
Shicong Feng1, Tianpeng Zheng2,3, Hao Hang4
1Graduate School of Education, Peking University, Beijing, China.
Medical Teacher
|November 21, 2025
Summary
Predicting item difficulty in medical education assessments is vital. Using the item stem and correct answer with XGBoost and domain-specific embeddings offers the best prediction accuracy.
Area of Science:
- Educational Measurement
- Machine Learning in Education
- Medical Education Assessment
Background:
- Accurate item difficulty prediction is essential for high-stakes educational assessments like medical licensing exams.
- Existing research shows inconsistent findings regarding influential modeling components, indicating a need for systematic investigation.
- Understanding these components is crucial for improving the design and administration of educational assessments.
Purpose of the Study:
- To systematically investigate key factors influencing item difficulty prediction performance in educational assessments.
- To identify the optimal combination of modeling components for accurate difficulty prediction.
- To provide insights for data-driven measurement practices in medical education.
Main Methods:
- The study explored the impact of model domain specificity, input content granularity (item stem, correct answer, distractors), embedding dimensionality, and machine learning regressor choice.
- 2815 Multiple-Choice Questions from the National Center for Health Professions Education Development were used to predict item difficulty.
- A range of embedding models and machine learning models were employed in the prediction task.
Main Results:
- XGBoost demonstrated superior performance among the machine learning regressors (Mean RMSE = 0.1779).
- A domain-specific embedding model (MedEmbed-small) consistently enhanced prediction accuracy (Mean RMSE = 0.1860).
- Inputting the item stem and correct answer yielded the best balance of predictive accuracy and model simplicity (RMSE = 0.1756).
Conclusions:
- The findings provide valuable insights for data-driven measurement practices such as Automated Item Calibration, Computerized Adaptive Testing, and Intelligent Tutoring Systems in medical education.
- Optimal feature sets for difficulty prediction are dependent on the specific item style.
- Future research should focus on extending these findings to predict the difficulty of multimodal test items.


