A Transfer Learning Method for Hate Speech Detection

Eva Šmuc, Goran Delač, Marin Šilić, Klemo Vladimir · 2023

In this work we explore the possibilities of using transfer learning techniques to enhance performance of hate speech detection models by relying on similar linguistic problems (e.g. toxic language detection). Multiple algorithms are trained for similar linguistic tasks on larger datasets, and the obtained models are used for getting predictions on the ETHOS dataset, which we chose as the target dataset of our work. The obtained predictions are used as sole or additional features in the subsequently performed experiments. Multiple algorithms are evaluated, including Logistic Regression, SVM, RidgeClassifier, Decision Tree, Random Forest, AdaBoost, GradBoost, Bagging. Furthermore, multiple textual representations are taken into account including Tf-Idf, Bert embeddings and BERT embeddings combined with the aforementioned additional features. Transformer-based models BERT and DistilBERT are introduced and fine-tuned on ETHOS dataset. All the obtained models are evaluated and the resulting performance metrics are compared to results obtained by the authors of the ETHOS dataset. In order to explore the remaining underlying issues, model-agnostic method LIME is used to obtain explanations for incorrect predictions.

Read the paper · More papers on PaperTik