Explaining Offensive Language Detection
Julian Risch, Robin Ruff, Ralf Krestel · LDV-Forum/Journal for language technology and computational linguistics · 2020
Machine learning approaches have proven to be on or even above human-level accuracy for the task of offensive language detection.In contrast to human experts, however, they often lack the capability of giving explanations for their decisions.This article compares four different approaches to make offensive language detection explainable: an interpretable machine learning model (naive Bayes), a model-agnostic explainability method (LIME), a model-based explainability method (LRP), and a self-explanatory model (LSTM with an attention mechanism).Three different classification methods: SVM, naive Bayes, and LSTM are paired with appropriate explanation methods.To this end, we investigate the trade-off between classification performance and explainability of the respective classifiers.We conclude that, with the appropriate explanation methods, the superior classification performance of more complex models is worth the initial lack of explainability. JLCL 2020 -Band 34 (1) -29-47Risch, Ruff, Krestel Machine-learned models, such as models that detect offensive language, should therefore be comprehensible.The field of Explainable AI (XAI) emerged to address this problem by making models interpretable and/or explainable.Explainable AI is a young and multidisciplinary research area, ranging from machine learning, data visualization, and human-computer interaction to psychology.Researchers distinguish between interpretability and explainability (Lipton, 2018).Interpretability means to convey a mental model of the algorithm to humans.In other words, if a model is interpretable, humans can grasp how its internals work.In contrast, explainability means to explain individual predictions of a model, rather than the full model itself.With an explainable model, humans can comprehend the calculation steps that lead from a particular input to a particular output.On the other hand, interpretability enables developers to understand a model's weaknesses and to improve on them.Some machine learning algorithms, such as decision trees, logistic regression, and naive Bayes, are interpretable by default.However, with an increasing number of features and sophisticated preprocessing, even these simple models lose their interpretability.More complex, non-linear models, such as neural networks and support vector machines (SVMs) with kernels, achieve better accuracy in some tasks but are not interpretable by default.So it might seem that there is a trade-off between accuracy and interpretability.Explainability is easier to achieve as it is sufficient to explain only single predictions of a model rather than the model itself.There are explainability methods that are specific to machine learning algorithms (model-based) and methods that can be applied to any model (model-agnostic, post-hoc).With many explained predictions of a black-box model, a human's mental model of the algorithm improves.Thereby, explainability can lead to interpretability.Recently, there is research that is contrary to post-hoc explanation methods.For example, Rudin (2019) states that the focus should be on creating inherently interpretable models rather than retrospectively explaining black-box models.With the General Data Protection Regulation (GDPR) 1 specifying the right to explanations, developing explainable AI systems is inevitable, and we expect the field of Explainable AI to grow in the future.Especially the highly complex neural networks with millions of parameters raise the bar for many natural language processing tasks significantly.At the same time, these models are non-interpretable black boxes.We see a need to make especially these most complex models explainable to ensure trust in them by humans.In our work, we train a naive Bayes classifier, an SVM, and recurrent neural network models on a dataset of toxic comments.We examine the explanation methods Layer-wise Relevance Propagation (LRP) and Local Interpretable Model-agnostic Explanations (LIME), but also attention layers.Thereby, our study covers a model-based method, a model-agnostic method, and a self-explanatory model.The naive Bayes classifier serves as a baseline.For the evaluation of the explanation methods, we use a word deletion task, the explanatory power index, and t-SNE projections of document vector representations.We discuss the results and find that the explainability methods LRP