Pre-processing Techniques for Performing Hotel Review Sentiment Analysis

Bharti B. Balande, Dinesh M. Kolte, Ramesh R. Manza, Suryakant S. Revate · 2023

In recent times, AI researchers have shown an increasing fascination with sentiment analysis, primarily due to its promising potential for profitable commercial uses. The initial stage of text mining in a Sentiment Analysis system involves crucial pre-processing steps. Given the frequent deviations from grammar and spelling norms in textual data, these steps become essential for data refinement. Moreover, prior to commencing the analysis it, pre-processing plays an essential part in standardizing the text. The main target of this investigation is to emphasize the importance of pre-processing techniques and showcase their capacity to enhance system accuracy. The study delves into diverse pre-processing techniques and compares their respective levels of accuracy. This comparative assessment aims to pinpoint the most effective strategies. In addition to a comprehensive evaluation of each method, the study delves into the rationale behind the observed accuracy enhancements. It's widely acknowledged that sentiments are often encapsulated in reviews, as individuals, particularly those who are enthusiastic or socially engaged, tend to articulate their opinions through this medium. Such reviews not only capture the viewpoints of their creators but also encapsulate their emotions. Given the unstructured nature of this textual content, effective handling necessitates the application of pre-processing techniques to prepare the data. Subsequently, the pre-processed data serves as a foundation for feature extraction. Approaches utilized to extract features encompass a range of methods, encompassing Bag of Words (BoW), Term Frequency - Inverse Document Frequency (TF-IDF), word embeddings, and Characteristics acquired through Natural Language Processing (NLP), such as word frequency and noun enumeration. These techniques constitute only a portion of the diverse methods employed.

Read the paper · More papers on PaperTik