Sentiment Categorization through Natural Language Processing

Jyovita Christi, Gayatri Jain · Journal of Emerging Technologies and Innovative Research · 2020

Sentiment is an attitude, thought, or judgement prompted by feeling. Sentiment Categorization, studies people’s sentiments towards certain units. Sentiment Analysis isn’t an unfamiliar term anymore. Today, smart phones, high speed and affordable Internet and various forums and social networks, have made it very common for people to give voice to their opinions. Therefore, a lot of textual data is available in various forms where people express their opinions. Analysing this data to know the underlying sentiment behind it has also become quite popular these days. Sentiment Categorization involves classifying text as positive or negative towards a certain target. The model proposed here, makes use of a hybrid approach of Natural Language Processing and Machine Learning to achieve an accuracy of 90% for an IMDB movie review dataset. To achieve this accuracy, the system uses methods like: Bag of words model, TF-IDF to calculate the relevance of each term in each sentence, regular expressions to remove punctuations and retain emojis by shifting them to the end of each sentence, tokenization and stemming to break the sentences into tokens and restore all words to their roots, stop word removal to remove words that do not bear any sentiment, and finally the logistic regression algorithm to perform the sentiment categorization into positive and negative.

Read the paper · More papers on PaperTik