Enhancing Online Safety: Automated Hate Speech Detection on Instagram with BERT-powered Model and Real-time Moderation

Rupali Gangarde, Moubani Ghosh, Ravi Thacker · 2024

This paper introduces an innovative hate speech detection model tailored for Instagram, utilizing the powerful Bidirectional Encoder Representations from Transformers (BERT) model. Through a meticulous integration of BERT with an elaborate data collection strategy and refined methodology, the study seeks to enhance precision and recall, mitigating false positives and negatives, thereby fostering a safer and more inclusive online environment on Instagram. Addressing the existing research gap, the model distinguishes itself by incorporating real-world context and addressing nuanced language nuances often missed by current hate speech detection models. The HateXplain dataset is employed for training and validation, complemented by a web scraping strategy for collecting real-time Instagram comments via the “Insc” extension. The BERT model undergoes fine-tuning to adapt to Instagram-specific language patterns. The comprehensive methodology spans data selection, training, and integration with Instagram, classification, evaluation, and real-time deployment. The anticipated contributions and findings include the model's superiority in accuracy and performance metrics, such as precision, recall, and F1-score, and its effectiveness in detecting diverse forms of hate speech while maintaining robustness. Comparative analysis showcases the advantages of BERT over LSTM models in terms of accuracy and contextual understanding. The paper suggests future directions, urging exploration of multilingual support, user profile integration for enhanced context, and experimentation with multi-modal hate speech detection incorporating images and videos. In summary, this paper offers a promising avenue for combating hate speech on Instagram, striving to contribute to the creation of a safer and more inclusive online community.

Read the paper · More papers on PaperTik