Improvisation in opinion mining using data preprocessing techniques based on consumer’s review

International Journal of Advanced Technology and Engineering Exploration · 2023

The consumer is one of the assets of every organization.In order to survive in this competitive marketplace, it is mandatory for every organization to assess the sentiments, expectations, and feedback of consumers.For this assessment, it is required to perform the sentiment analysis of data collected from various internet sources.The data collected from various internet sources contain various types of anomalies such as stop words, hypertext markup language (HTML) tags, misspellings, abbreviations, special characters, uniform resource locator (URL), etc. Due to this, it becomes difficult for both humans and machines to get the exact meaning of the sentence. *Author for correspondenceMoreover, the presence of this unwanted noise in data increases the dimensions because each word in the text is considered a separate feature and it becomes challenging for classifiers to classify this noisy data.So, to get accurate sentiments it is mandatory to remove this unwanted noise from the data.In previous studies, various preprocessing techniques were used for different types of datasets and all these techniques were evaluated to find the best preprocessing techniques using various classifiers such as support vector machine (SVM), logistic regression (LR), decision tree (DT), random forest (RF), Naïve bayes (NB).Some of the basic preprocessing approaches like removal of stop words, punctuation, tokenization, spell correction [15] stemming/lemmatization were used in various pieces of research.Similarly, various

Read the paper · More papers on PaperTik