ULMFiT at GermEval-2018: A Deep Neural Language Model for the Classification of Hate Speech in German Tweets
Kristian M. Rother, Achim Rettberg · ISBN · 2018
This paper describes the entry hshl coarse 1.txt for Task I (Binary Classification) of the Germeval Task 2018 - Shared Task on the Identification of Offensive Language. For this task, German tweets were classified as either offensive or non-offensive. The entry employs a task-specific classifier built on top of a medium-specific language model which is built on top of a universal language model. The approach uses a deep recurrent neural network, specifically the AWD-LSTM architecture. The universal language model was trained on 100 million unlabeled articles from the German Wikipedia and the medium-specific language model was trained on 303,256 unlabeled tweets. The classifier was trained on the labeled tweets that were provided by the organizers of the shared task.