Multi-Class Emotion Classification on Tamil and Tulu Code-Mixed Text

N Prabhu Ram, T. Meeradevi, P. Sendhuraharish, S Yogesh, C VasanthaKumar · 2024

Sentiment analysis, which is frequently used to gauge public sentiment in reviews and social media posts, is the technique of analyzing and interpreting feelings and viewpoints in text data. The surge in social media users and the proliferation of user-generated code-mixed content, including reviews, comments, and posts, has spurred a growing demand for efficient tools that can adeptly assess this content to identify underlying sentiments. Nonetheless, analyzing sentiments in social media text proves to be a challenging task due to the intricate and often mixed linguistic nature of the content. To tackle this challenge, the mBERT model, yielded impressive results. Experiments with a transfer learning model found that removing high-frequency tokens from the mixed feeling class, compared to all other classes, had a detrimental effect on the model's performance. Throughout the study, carefully examined the repercussions of token removal on the model's overall performance. Specifically, focused on assessing whether the elimination of these high-frequency tokens from various classes significantly affected the model's proficiency in identifying and classifying emotions or feelings within the text data. These findings emphasize the crucial need to carefully evaluate the impact of token removal strategies on a model's capability to detect emotional content. The mBERT model demonstrated remarkable performance with a high accuracy of 0.78 and a strong F1-score of 0.78.

Read the paper · More papers on PaperTik