A Multilingual BERT-Based Framework for Robust Online Hate Speech Detection

Yash Shukla, Rajkumar Vamanrao Panchal, Tanmay Nigade, Suyash Khodade, Prathamesh Pimpalkar · 2024

Hate speech poses a significant threat to societal harmony and safety, often inciting violence and discrimination. Recent real-time incidents, such as the communal unrest in various parts of the world triggered by inflammatory online content, demonstrates the need for effective hate speech detection mechanisms. In this paper, we present a comprehensive hate speech detection model utilizing the Bidirectional Encoder Representations from Transformers (BERT) algorithm, leveraging advanced Deep Learning and Natural Language Processing (NLP) techniques. Our model is designed to address the challenges of detecting hate speech within the Hinglish language, a code-mixed variant combining Hindi and English. We used a pre-trained BERT model and tailored that model to function in real-time scenarios, offering robust and in-depth analysis capabilities. We adopted a multilingual and multimodal approach, enabling our system to detect hate speech. This task is accomplished across various content formats, including text, audio, video, images, GIFs, and YouTube video comments. Experimental results demonstrate that our model achieves an accuracy of 0.88. This underlines the model’s efficacy in detecting hate speech, showcasing its potential for application in diverse real-world contexts.

Read the paper · More papers on PaperTik