Incepto@DravidianLangTech 2025: Detecting Abusive Tamil and Malayalam Text Targeting Women on YouTube
Luxshan Thavarasa, Sivasuthan Sukumar, Jubeerathan Thevakumar · 2025
This study introduces a novel multilingual model designed to effectively address the challenges of detecting abusive content in lowresource, code-mixed languages, where limited data availability and the interplay of mixed languages, leading to complex linguistic phenomena, create significant hurdles in developing robust machine learning models.By leveraging transfer learning techniques and employing multi-head attention mechanisms, our model demonstrates impressive performance in detecting abusive content in both Tamil and Malayalam datasets.On the Tamil dataset, our team achieved a macro F1 score of 0.7864, while for the Malayalam dataset, a macro F1 score of 0.7058 was attained.These results highlight the effectiveness of our multilingual approach, delivering strong performance in Tamil and competitive results in Malayalam.