A Scalable Model for Text Mining using Sentiment Analysis on PySpark

K S Vishal, Srivageesh K Srinidhi, U Someswara Shashank, Sankepalli Sai Samhitha, Manju Venugopalan · 2024

Sentiment analysis, an essential tool for deciphering public sentiment from vast internet data, offers valuable insights to businesses and policymakers. Its adoption is driven by the ability to interpret emotions, providing guidance for product refinement and decision-making. This tool is crucial for understanding user sentiments, empowering businesses to improve customer experiences, shape brand perception, and refine marketing strategies. Organizations benefit from sentiment analysis by leveraging customer insights for informed decision-making and product enhancement, thereby safeguarding positive relationships and brand reputation. The proposed system utilizes both formal and informal datasets and works on models like Logistic Regression, Naive bayes and SVM. Additionally, feature extraction techniques such as hashing and TF-IDF enhance data representation, contributing to the accuracy and F1-score of the employed models. The proposed system makes use of the PySpark platform to elevate the efficiency.Among the three models utilized, SVM achieved the highest accuracy and F1-score for the Twitter dataset, with an accuracy of 0.76 and an F1-score of 0.76. Similarly, for the Amazon dataset, Logistic Regression exhibited the highest accuracy and F1-score, recording values of 0.89 for both metrics. For comparison and experimental evaluation, the models were executed on a merged dataset comprising both Twitter and Amazon data. In this merged dataset, Logistic Regression outperformed the other models, achieving an accuracy and F1-score of 0.83 for both metrics.

Read the paper · More papers on PaperTik