Detecting Hate Speech for Hindi-English Code-Mix Text Data Using Dual Contrastive Learning
Amit Sharma, Rajni Bhalla · Procedia Computer Science · 2025
A bigger problem in today’s world is hate speech on internet, especially as social media sites like Facebook and Twitter rapidly grow. It includes any speech or writing that makes fun of or calls for violence against people or groups because of things that make them unique, like race, religion, gender, or sexual orientation. This paper looks at how to find hate speech in Hindi-English Code-mixed text data using both Traditional machine learning (ML) models and Dual Contrastive Learning (DCL) method. Word embeddings, TF-IDF, and pre-trained SentBERT techniques were used along with other methods to prepare and examine the data. Traditional ML models like KNN, DT, and LR are tested in the study. These models get accuracy rates of 73.63%, 79.78%, and 85.50%, respectively. The proposed DCL model does better than these, as its accuracy rate is 86% and its F1-Score is 91%. The results show that pre-trained SentBERT could help better as extracting the feature in code-mixed data. The accuracy, on the other hand, shows that there is room for improvement. Adding more advanced deep learning and natural language processing methods could make the model even better.