Evaluating the Consistency and Reliability of Attribution Methods in Automated Short Answer Grading (ASAG) Systems: Toward an Explainable Scoring System
Wallace Nascimento Pinto, Jinnie Shin · Journal of Educational Measurement · 2025
Abstract In recent years, the application of explainability techniques to automated essay scoring and automated short‐answer grading (ASAG) models, particularly those based on transformer architectures, has gained significant attention. However, the reliability and consistency of these techniques remain underexplored. This study systematically investigates the use of attribution scores in ASAG systems, focusing on their consistency in reflecting model decisions. Specifically, we examined how attribution scores generated by different methods—namely Local Interpretable Model‐agnostic Explanations (LIME), Integrated Gradients (IG), Hierarchical Explanation via Divisive Generation (HEDGE), and Leave‐One‐Out (LOO)—compare in their consistency and ability to illustrate the decision‐making processes of transformer‐based scoring systems trained on a publicly available response dataset. Additionally, we analyzed how attribution scores varied across different scoring categories in a polytomously scored response dataset and across two transformer‐based scoring model architectures: Bidirectional Encoder Representations from Transformers (BERT) and Decoding‐enhanced BERT with Disentangled Attention (DeBERTa‐v2). Our findings highlight the challenges in evaluating explainability metrics, with important implications for both high‐stakes and formative assessment contexts. This study contributes to the development of more reliable and transparent ASAG systems.