Sentiment Analysis on Comments in Bengali Language Using Text Mining & Machine Learning Approach

Priya Das, Nasrin Sultana · 2022 IEEE 7th International conference for Convergence in Technology (I2CT) · 2022

In the modern era, people are highly engaged in the virtual world to express their feelings and opinions. Social platforms and news portals are the precious media for expressing views on any sector nowadays. Thousands of information are being added to social media sites every second. During this research, some specks of data were collected and analyzed for extracting public sentiment. Sentiment Analysis is the process to identify computationally and categorize opinions expressed in a piece of text significantly to determine the author's attitude towards a specific topic. It is one of the most renowned research areas in Data Science. There is much research done on sentiment analysis in various sectors for the English language and increasing. Still, works for the Bengali language are limited to Bangla corpus and Bangla micro-blogging. So, this research has aimed to apply Sentiment Analysis to the Bengali language in which the public can express their opinions and emotions in their endemic words. However, it is not easy to apply it to the Bengali language because of the complex grammatical structure of Bangla. This paper represents the method of preprocessing the dataset using the NLP and applying Machine Learning approaches to classify test data. The process starts with removing unnecessary words and then using the TF-IDF vectorizer for feature extraction and applying cosine similarities, LSA and SVD for feature selection to form a meaningful and optimized feature vector to undertake the experiments. Then an analysis was performed to compare different machine learning approaches to the extracted information. Finally, the proposed model categorizes the document into five classes, namely negative, near negative, neutral, near positive, and positive, with profound accuracy, which might further be applied on similar types of datasets.

Read the paper · More papers on PaperTik