LSTM Neural Networks for Transfer Learning in Online Moderation of Abuse Context
Avi Bleiweiss · 2019
Recently, the impact of offensive language and derogatory speech to online discourse, motivated social media platforms to research effective moderation tools that safeguard internet access. However, automatically distilling and flagging inappropriate conversations for abuse remains a difficult and time consuming task. In this work, we propose an LSTM based neural model that transfers learning from a platform domain with a relatively large dataset to a domain much resource constraint, and improves the target performance of classifying toxic comments. Our model is pretrained on personal attack comments retrieved from a subset of discussions on Wikipedia, and tested to identify hate speech on annotated Twitter tweets. We achieved an F1 measure of 0.77, approaching performance of the in-domain model and outperforming out-domain baseline by about nine percentage points, without counseling the provided labels.