Sentiment Analysis in Code-Mixed Telugu-English Text (CMTET) Using Deep Learning Techniques

Nachiketh Velekkat Suraj, Anil Kumar Manda, Vinay Raj · 2024

In social media, individuals often communicate by combining two or more languages leading to the emergence of Code-Mixed data. Analyzing sentiment in Code-Mixed Telugu-English Text (CMTET) provides unique issues owing to its unstructured nature, which includes informal language, transliterations, and spelling mistakes. To overcome these challenges, this study explores various deep learning models such as RNN, GRU, BiLSTM, CNN and BERT for performing sentiment analysis in CMTET by leveraging on an annotated dataset which contains social media reviews. From the results, it is observed that BERT model combined with novel unsupervised data normalization technique yielded an accuracy of 79.72%. The proposed data normalization technique demonstrated in this paper holds potential for extension to various NLP tasks involving CMTET.

Read the paper · More papers on PaperTik