Enhancing Hate Speech Detection in Mixed-Language Texts: A Comparative Study of BLOOM and XLM-RoBERTa Models

Arya Dimas Wicaksana, Kevin Sorensen, Farrel Dinarta · 2025

Hate speech detection is essential in combating online toxicity, particularly in mixed-language or code-switched texts prevalent on social media. Traditional natural language processing (NLP) models often struggle with these complex linguistic structures due to blending multiple languages. This paper investigates the effectiveness of BLOOM (BigScience Large Open-science Open-access Multilingual language model) and XLM-RoBERTa, two powerful multilingual models, in addressing these challenges. BLOOM's extensive pre-training across diverse languages and XLM-RoBERTa's robust capabilities allow a nuanced understanding of context in mixed-language environments. We fine-tune both models on an English and Indonesian text dataset containing instances of mixed-language hate speech and evaluate their performance against state-of-the-art benchmarks. Our findings highlight the effectiveness of these models in recognizing hate speech in mixed-language scenarios, with the fine-tuned BLOOM (bloom-560m) performing better than XLM-RoBERTa (xlm-roberta-base).

Read the paper · More papers on PaperTik