Evaluating Explanation Fidelity for Transformer-Based Software Models
Anonymous · Zenodo (CERN European Organization for Nuclear Research) · 2025
Foundation models such as CodeBERT, GraphCodeBERT, and CodeT5 have transformed software engineering by enabling powerful tools for defect prediction, bug localization, and code summarization. Yet their complex, opaque decision-making often leaves developers questioning why a model makes a particular prediction, hindering trust and adoption. Motivated by this gap, our work focuses on building faithful and developer-trustworthy explanations for these transformer-based models. We introduce CoScoreX (Contextual Score Explanation), a hybrid framework that combines contextual embeddings with contrastive reasoning to generate token-level explanations that are both accurate to the model’s behavior and meaningful to developers. We evaluate CoScoreX alongside six widely used explainers like SHAP, LIME, Integrated Gradients, Raw Attention, Attention Rollout, and Grad-Attention Rollout, across four foundation models (three for defect prediction and one for sentiment analysis) using Comprehensiveness and Sufficiency metrics with statistical validation (Wilcoxon and Friedman tests). The results show that CoScoreX consistently provides the most balanced and stable explanations across models and domains. In this work, we first present a unified fidelity benchmarking pipeline for transformer-based code models. Secondly, we introduce a contextual–contrastive explanation framework that enhances interpretability. Thirdly, we demonstrate strong cross-model generalization of explanation fidelity. Finally, we offer actionable insights aligned with the AI's vision of trustworthy and developer-centered foundation models, paving the way for future extensions through causal reasoning and human-in-the-loop evaluation.