Interactively Providing Explanations for Transformer Language Models
Felix Friedrich, Patrick Schramowski, Christopher Tauchmann, Kristian Kersting · Frontiers in artificial intelligence and applications · 2022
Transformer language models (LMs) are state of the art in a multitude of NLP tasks.Despite these successes, their opaqueness remains problematic, especially as the training data might be unfiltered and contain biases.As a result, ethical concerns about these models arise, which can have a substantial negative impact on society as they get increasingly integrated into our lives [1].Therefore, it is not surprising that a growing body of work aims to provide interpretability and explainability to black-box LMs [2]: Recent evaluations of saliency or attribution methods [3,4] find that, while intriguing, different methods assign importance to different inputs for the same outputs, thus encouraging misinterpretation and reporting bias [5,6].Moreover, these methods primarily focus on post-hoc explanations of (sometimes spurious) input-output correlations.Instead, we emphasize using (interactive) prototype networks directly incorporated into the model architecture and hence explain the reasoning behind the network's decisions.