Code-Mixed Romanized Hindi Hate Speech Identification: Leveraging BERT Embeddings and Particle Swarm Optimization
Shubham Shukla, Sushama Nagpal, Sangeeta Sabharwal · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
The volume of hate speeches and the number of user-generated materials are steadily rising, notably on social media networks. This trend can be seen across the internet. Therefore, it is necessary to recognize this kind of offensive content and remove it to maintain the cleanness of the platform. In spite of the fact that pertinent research was conducted separately for detecting hate speeches and social media code-mixed texts, purpose of this work is to identify hate speeches from social media code-mixed text. In the presented research, experiments for detecting hate speech using Bidirectional Encoder Representations from Transformers (BERT) architecture are carried out with the accessible code-mixed dataset, and optimization is carried out using the Particle Swarm Optimization algorithm (PSO). The results of our proposed methodology showcased notable improvements in the performance over standard BERT model while achieving an accuracy of 95.37% and an F1-score of 95.30%. Further, testing on an additional dataset confirms the generalizability of our approach, maintaining a high accuracy of 94.5%. These findings validate the effectiveness of PSO-driven optimization for hate speech detection in low-resource, code-mixed linguistic settings. To offer a comprehensive assessment of our approach, a comparative analysis with several classifiers, including Naive Bayes, Random Forest, XG Boost, CatBoost, KNN, Decision Tree, Adaboost, SVM, and LSTM has also been done.