Sentiment Analysis for Twitter Data in the Hindi Language

Anjum Madan, Udayan Ghose · 2021

In the past decade, mining the opinions for information extraction and retrieval has been one of the major fields of study. Sentiment Analysis (SA) can be explained as the process of identifying the polarity of a given content for gaining an insight into the hidden information stored in a text or a piece of document. Today, various social networking sites serve as a platform for expression. Twitter, is one such microblogging application that is tremendously used by various users for expressing the true intent about a particular entity in the form of tweets. The ultimate goal of analyzing the sentiments is to mine the useful content from these opinion sources. The Natural Language Processing (NLP) technique has been used for the initial processing of tweets. There are broadly two methods available for further analyzing the opinionated data: Lexicon Based Approach (LBA) and Machine Learning Approach (MLA) based on supervised learning. The Supervised method has a disadvantage of its dependency on the quality and amount of training data. The LBA technique uses an enhanced dictionary i.e. Hindi SentiWordNet as a resource and Hybrid Based Approach (HBA) which combines the LBA and MLA for classifying movie tweets as either positive or negative. The English language has majorly dominated in this area of research. We, on the other hand, have gathered movie tweets in the Hindi language for Sentiment Analysis of Twitter data and led a comparative evaluation for both the techniques: LBA and HBA.

Read the paper · More papers on PaperTik