Hate Speech Detection in Low-Resource Languages Hindi, Gujarati, Marathi, and Sinhala

Enduri Jahnavi, Animesh Chaturvedi · 2025

Sentiment analysis in low-resource languages poses unique challenges due to limited linguistic resources and labeled datasets. This paper focuses on low-resource languages spoken in India and Sri Lanka: Hindi, Gujarati, Marathi, and Sin-hala. The study explores machine learning and deep learning models, including BERT, ensemble learning, and various deep learning architectures, to address the challenges of sentiment analysis in these languages. The effect of stacking the same and different deep learning models on top of each other and comparing them based on accuracy and F1-score is studied. We proposed an ensemble BERT+CNN+LSTM model for detecting hate speech in these languages. The study also highlights the significance of thoughtful model selection and parameter tuning for optimal sentiment classification in low-resource languages through various experiments on model architectures and varying hyper parameters. Our results emphasise the importance of understanding sentiment for public opinion, social trends, and communication strategies.

Read the paper · More papers on PaperTik