Ensemble of Transformer Based Approach for Hate Speech Detection on Twitter Data

Kartik Singh, Meenakshi Tripathi, Basant Agarwal, Abhay Kumar Sain · 2023

Social media platforms like Twitter are popular for sharing opinions and ideas, but hate speech is a growing concern. Hate speech can harm individuals and communities, leading to discrimination and violence. Detecting and preventing hate speech on social media is now crucial. Different models such as BERT, CNN, LSTM, and XLM-ROBERTA have shown promising results in identifying hate speech, but each has its advantages and limitations. For example, BERT is good at capturing semantic meaning but not local patterns, while CNN is good at detecting local patterns but not context. LSTM can capture temporal dynamics but may struggle with long-term dependencies. XLM-ROBERTA can handle multilingual text but may not perform as well on certain types of hate speech. The BERT, BERT-CNN, and XLM-ROBERTA models have consistent outcomes but low precision in identifying hate speech. The BERT-LSTM model has the highest precision but lower overall performance. In this paper, we present ensembling approaches that combines the strengths of multiple models, more specifically, the proposed ensemble model relying on predictions from BERT, BERT-CNN, and XLM-ROBERTA, with BERT-LSTM used as a tiebreaker for hate speech classification. The ensemble learning methods described in this paper outperformed various deep learning models for identifying hate speech.

Read the paper · More papers on PaperTik