Sentiment Analysis using Transformer Based Pre-Trained Models for the Hindi Language

Akshat Verma, Shivam Walbe, Ishwar Wani, Ritesh Wankhede, Radha Thakare, Sanika S Patankar · 2022

India has one of the largest user bases for Internet-based services. In 2021, India had 624 million internet users, the second largest in the world11https://datareportal.com/reports/digital-2021-india. With so many internet users, each generating lots of textual data, having tools to analyze the data can be very helpful to a wide variety of people including researchers, marketers, and product managers. Research regarding Sentiment Analysis in English is plentiful, but we need a different method to perform the same in Hindi, one of the most popular languages of the Indian subcontinent. In this paper, we use Transformer-based pre-trained models on Hindi Sentiment Analysis tasks. The sentiment analysis task is done on a dataset that has Hindi text and its corresponding sentiment as Positive, Negative, and Neutral. The Hindi text contains sentences collected from various sites. The sentences primarily contain product and movie reviews. The sentiment analysis is done using five different transformer-based models, out of which a few have been trained for multiple languages while the others have been fine-tuned specifically for the Hindi language. We also compare the performance of a few different multilingual models on sentiment analysis tasks. Out of all the models compared, we get the best accuracy of 82% from the Hindi Microsoft Multilingual-MiniLM-L12-H384 model.

Read the paper · More papers on PaperTik