Towards Effective Detection of Inappropriate Language in Social Media Context Using BERT

Sankranthi Varshitha, P G Suthiksha, S Natarajan · 2024

In the fast-growing phase of social media, detecting the offensive language has become more crucial, since people are using different languages and cultural expressions. This paper represents a new way to handle the problems faced to detecting offensive language in different languages and understanding in which context it has been used. We have used BERT that is Bidirectional Encoding Representation from Transformers, a powerful learning tool where it will handle multiple datasets such as Tamil-English, Malayalam-English. This approach is more accurate and efficient when compared with the traditional models such as Random Forest and Decision Tree. This tool is designed in such a way that it enables easier understanding for both the native and non-native speakers in order to understand the foul or offensive content which is in context. Our model addresses the challenge of detecting offensive language in multilingual data and focuses more on understanding the sarcasm, humor and slang which machines predict inaccurately.

Read the paper · More papers on PaperTik