Comparative Analysis of Different BERT-Based Machine Learning Models for Hostile Hindi Content Detection

Angana Chakraborty, Subhankar Joardar, Arif Ahmed Sekh · 2023

Platforms on social media increasingly offer hostile content. As a result, exact hostile post detection is now necessary in order to launch the proper response. Language comprehension now faces new challenges due to the increasingly hostile content over different electronic medias. Language barriers increase the difficulty. Even though there have been many research conducted in English, regional languages haven’t made much progress as the necessary tools and datasets aren’t currently accessible. In this research, contextual embedding method based on Bidirectional Encoder Representations from Transformers (BERT) is combined with Support Vector Machine (SVM) to distinguish Hindi Hostile or Non-Hostile posts on social networking sites by employing Constraint 2021 Hindi Dataset. Various cutting-edge BERT-based methods are also compared and analyzed with machine learning methods like Logistic Regression (LR), Support Vector Machine (SVM), Random Forest (RF), Multilayer Perceptron (MLP) in this research work. For the four hostile subclasses (Defamation, Fake, Hate, and Offensive), it is discovered that the performance of the model we propose (Indic-BERT+SVM) surpasses the baseline model with F1-Score of 96.49 for binary (hostile/non-hostile) classification task and F1-Scores of 44.63, 76.00, 56.93, 61.80 respectively for multi-class, multi-label (defamation, fake, hate, and offensive) classification tasks.

Read the paper · More papers on PaperTik