Evaluating explainability in language classification models: A unified framework incorporating feature attribution methods and key factors affecting faithfulness

Tahereh Dehdarirad · Data and Information Management · 2025

This paper presents a unified framework for evaluating explainability methods in language classification models, integrating feature attribution and interaction approaches while considering key factors impacting faithfulness: model architecture, dataset characteristics, and evidence type. By comparing classical (Logistic Regression and Random Forest) and transformer models (RoBERTa and DistilBERT), the faithfulness of SHAP, LIME, and Integrated Gradients (IG) across positive, negative, and all evidence types were examined. In classical models, SHAP and LIME generally provide faithful explanations for positive and all evidence types, with SHAP and Random Forest best handling negative evidence. For transformer models, the faithfulness of LIME and SHAP varies by model, dataset, and evidence type. LIME performs consistently well in complex models like RoBERTa and DistilBERT, while SHAP excels with positive evidence across datasets and is most effective in RoBERTa for negative evidence. For all evidence, SHAP also shows broader applicability across evidence types, whereas LIME is suited to specific datasets, such as Brexit, especially in RoBERTa and DistilBERT. For longer texts, IG and SHAP outperform LIME, with SHAP excelling in complex architectures like RoBERTa. When using IG, RoBERTa provides slightly more faithful explanations than DistilBERT for positive evidence, though only DistilBERT aligns with expected trends for negative evidence. Feature interaction analyses using Shapley Taylor Interaction (STI) and Archipelago reveal that RoBERTa consistently provides more cohesive explanations than DistilBERT across datasets, especially with Archipelago. STI-based models produce more interpretable, human-relevant phrases, achieving higher relevance ratings, especially when evaluated within full text contexts. • SHAP and LIME explain positive evidence well in classical models. • SHAP is best for negative evidence in Random Forest and RoBERTa. • They vary by model, dataset, and evidence in transformer architectures. • In longer texts, SHAP outperforms LIME, especially in RoBERTa. • STI with transformers yields feature phrases matching human judgments.

Read the paper · More papers on PaperTik