Bias explained: Generation of high-quality natural language explanations for classification decisions
Tim Menzner, Jochen L. Leidner · Expert Systems with Applications · 2025
It is desirable for machine learning classifiers to provide human-understandable explanations that justify decisions (“explainable artificial intelligence”, XAI). However, state-of-the-art neural network-based models including those used by large language models (LLMs) are intrinsically black-box approaches. To remedy this in the context of a classification model for news bias detection and sub-categorization, we present and evaluate a method for natural language explanation generation. We conduct a study in English with human subjects who were asked to rate machine-generated explanations for sentence-level bias classification decisions for news articles, where our method and a range of alternatives were compared. Our findings suggest that (1) explanations generated by LLMs significantly enhance the comprehensibility of classification decisions and contribute to a greater understanding of how news bias manifests in reporting and (2) one of our methods substantially outperforms all other methods, although we noticed two distinct preferential sub-groups. To the best of our knowledge, ours is the first study to evaluate LLM-generated explanations for news bias classification decisions by directly comparing zero-shot, fine-tuned, and traditional Natural Language Processing (NLP) approaches using feedback from regular users rather than trained experts or synthetic metrics.