Medical exam question difficulty prediction: An analysis of embedding representations, machine-learning approaches, and input feature impact
Shicong Feng, Tianpeng Zheng, Hao Hang, Jiayi Liu, Zhehan Jiang · Medical Teacher · 2025
Introduction Item difficulty prediction is crucial for planning and administrating educational assessments, especially those with high-stakes such as medical licensing examinations. The inconsistent findings across existing studies, however, highlight a critical gap in understanding which modeling components are most influential. This research addresses this gap by systematically investigating several key factors hypothesized to affect prediction performance.Methods This study explored the impact of: (1) model domain specificity, (2) input content granularity (e.g. item stem, correct answer, and distractors), (3) embedding dimensionality, and (4) the choice of the machine learning regressor. By selecting a range of embedding models and a series of Machine Learning models to predict the difficulty of 2815 Multiple-Choice Questions sourced from the National Center for Health Professions Education Development.Results Analyses revealed that XGBoost outperformed other counterparts (Mean RMSE = 0.1779), and the use of a domain-specific MedEmbed-small embedding model consistently improved prediction accuracy (Mean RMSE = 0.1860). Notably, using the item stem and the correct answer as input features achieved the best trade-off between predictive accuracy and model parsimony (RMSE = 0.1756).Discussion These findings offer valuable insights for data-driven measurement practices including Automated Item Calibration, Computerized Adaptive Testing, and Intelligent Tutoring Systems in medical education. Furthermore, this study revealed that the optimal feature set for difficulty prediction is contingent on the item style. Future research should extend this line of inquiry to the difficulty prediction of Multimodal test items. Practice pointsThe choice of machine learning algorithm, particularly XGBoost, is the most critical factor for accurate item difficulty prediction.Domain-specific embeddings greatly improve predictive performance over general-purpose models in medical education contexts.Using the item stem and correct answer as input features provides an optimal balance between prediction accuracy and model parsimony.A computationally efficient, high-performing model is achievable without relying on larger, more complex deep learning architectures.These findings provide a practical, low-cost framework for developing AI-assisted assessment tools like automated test assembly.