Navigating Emotion in Code-Mixed Languages: Performance of ML and DL Models on Hindi-English Text

Rutal S. Mahajan, Anjali S. More, Unnati S. Shah · Procedia Computer Science · 2025

With the rapid growth of internet usage, text data offers unprecedented insights into user behavior. This study addresses sentiment analysis and emotion recognition in Hindi-English code-mixed text—an underexplored area despite its significance in a multilingual country like India. Existing datasets primarily focus on single-label emotions, overlooking the complexity of mixed emotional states commonly observed in real-world scenarios. To address this gap, a novel multilabel dataset of 13,000 YouTube comments was developed, with 5,720 samples annotated to capture both single and mixed emotions. This dataset facilitates both coarse-grained (positive/negative) and fine-grained emotion recognition tasks, providing a versatile resource for research. The study evaluates several machine learning and deep learning models, including Logistic Regression, SVM, Naive Bayes, Random Forest, LSTM, BiLSTM, GRU, and CNN. Among these, SVM achieves the highest accuracy of 81.40% with an F1 score of 80.39%. In contrast, deep learning models like LSTM and GRU show lower performance, with an accuracy of 69.57% and an F1 score of 57.09%, indicating difficulties in detecting nuanced emotions such as joy and trust. This work offers a foundation for developing emotion recognition models that can effectively handle both single and mixed emotional states, reflecting real-world complexities. Future efforts will explore advanced architectures and extend this approach to other code-mixed languages.

Read the paper · More papers on PaperTik