Multilingual Offensive Language Detection In Social Media Content Using BERT-Base-Multilingual-Cased Model
P.S Nandhini, R. Karunamoorthi, P Mariappan, Revathi S A · 2024
The rise of social media and online communication has fostered global dialogue, transcending linguistic and geographic barriers. However, it has also ushered in the use of offensive language, posing challenges to safe digital spaces. Addressing this in a multilingual context is crucial for a more inclusive online environment. This paper tackles the classification of offensive and non-offensive comments, delving into the complexities of multilingual offensive language detection. Utilizing a joint-multilingual approach-based BERT model, comments are classified across various languages. Experiments on a dataset featuring English, Hindi, Telugu, Malayalam, Kannada, Greek, and Russian comments yield promising results: 92.47% of accuracy and a confidence score of 51.60%.