Evaluating Sentiment Analysis Algorithms with Varied Datasets: A Machine Learning Approach
S Indhumathi, F. Mary Harin Fernandez · 2024
Sentiment analysis is essential for comprehending public opinion on a variety of platforms and offers insightful information to researchers, businesses, and legislators. This study evaluates five machine learning algorithms-Support Vector Machine (SVM), K-Nearest Neighbors (KNN), Random Forest (RF), Logistic Regression (LR), Naive Bayes (NB) -applied to two distinct datasets: IMDB movie reviews and Twitter Airline sentiments. To provide a thorough performance comparison, each model is evaluated using recall, accuracy, precision, and F1 score. SVM achieves 91.02% accuracy on the Twitter dataset, making it the best performance; on the IMDB dataset, Logistic Regression wins with 87.39% accuracy. Naive Bayes demonstrates high precision, particularly on the Twitter dataset, though its recall score is relatively lower, indicating it is more conservative in capturing true positives. Random Forest maintains a balanced performance across all metrics, whereas KNN underperforms, particularly in handling the complexities of the IMDB dataset with a lower accuracy of 71.08%. The analysis reveals that dataset-specific factors significantly influence the performance of each algorithm, highlighting the importance of data characteristics in determining the most suitable model. While SVM excels in high-dimensional data like Twitter, Logistic Regression proves more effective with the IMDB dataset's structured text. This reinforces the need for model selection in sentiment analysis, depending on the nuances of each dataset. By providing a detailed evaluation across different models and metrics, this study contributes to optimizing sentiment analysis techniques for diverse applications, ensuring more robust and accurate predictions.