Aspect-Based Sentiment Analysis by Leveraging Machine Learning Techniques*

Yadavalli Uday Shankar, Somya Ranjan Sahoo, Saroj Kumar Panigrahy, Medeswara Rao Kondamudi · 2024

Based Sentiment Analysis (ABSA) is a task in Natural Language Processing that seeks to identify and extract the sentiment associated to specific parts or aspects of a product or service. ABSA often entails a sequential procedure that commences with the identification of the specific components or characteristics of the product or service that are being addressed in the text. Next, sentiment analysis is conducted to give a sentiment polarity to each aspect, taking into account the context of the sentence or document. Ultimately, the outcomes are combined to generate a comprehensive emotion for each individual feature. The technique entails training machine learning models to categorize the sentiment of text as either positive, negative, or neutral. Initially, we convert textual data using the Term Frequency-Inverse Document Frequency (TF-IDF) technique, which assigns weights to words depending on their significance within a collection of documents. This highlights the use of useful terminology. Subsequently, the TF-IDF features are inputted into the machine learning models. SVM determine an optimal hyperplane to effectively distinguish sentiment classes, whereas Logistic Regression computes the likelihood of a text being assigned to a certain sentiment class. Random Forest, on the other hand, constructs multiple decision trees and aggregates their results to enhance the accuracy and robustness of sentiment analysis. A series of comprehensive experiments were carried out on covid vaccinations dataset. The results indicate that the Logistic Regression model demonstrates exceptional performance in both aspect extraction and sentiment classification. The sentiment expressed on Twitter can exhibit an imbalance, with a prevalence of either positive or negative tweets contingent upon the subject matter. This can have an impact on the training process. Utilizing techniques such as oversampling or undersampling may be required for the minority class. This study examines the efficacy of machine learning algorithms in a particular classification challenge. The performance of Support Vector Machine (SVM), Logistic Regression (LR), and Random Forest was examined. The results demonstrate that Logistic Regression outperformed the SVM and Random Forest in terms of accuracy, achieving a rate of 92.87% compared to 91% and 87%, respectively. This suggests that Logistic Regression is a more appropriate choice for this classification task, given its superior accuracy and overall nerformance.

Read the paper · More papers on PaperTik