MSM_CUET@DravidianLangTech 2025: XLM-BERT and MuRIL Based Transformer Models for Detection of Abusive Tamil and Malayalam Text Targeting Women on Social Media
Md Mizanur Rahman, Srijita Dhar, Md. Mehedi Hasan, Hasan Murad · 2025
Social media has evolved into an excellent platform for presenting ideas, viewpoints, and experiences in modern society.But this large domain has also brought some alarming problems including internet misuse.Targeted specifically at certain groups like women, abusive language is pervasive on social media.The task is always difficult to detect abusive text for low-resource languages like Tamil, Malayalam, and other Dravidian languages.This paper presents a novel approach to detecting abusive Tamil and Malayalam texts targeting social media.A shared task on 'Abusive Tamil and Malayalam Text Targeting Women on Social Media Detection' has been organized by DravidianLangTech at NAACL-2025.We have implemented our model with different transformer-based models like XLM-R, MuRIL, IndicBERT, and mBERT transformers and the Ensemble method with SVM and Random Forest for training.We selected XLM-RoBERT for Tamil text and MuRIL for Malayalam text due to their superior performance compared to other models.After developing our model, we tested and evaluated it on the DravidianLangTech@NAACL 2025 shared task dataset.We found that XLM-R achieved the highest F1 score of 0.7873 on the test set, ranking 2 nd among all participants for abusive Tamil text detections.In contrast, MuRIL had the highest F1 score of 0.6812 for abusive Malayalam text detections, ranking 10 th among all participants.