Fine-Grained Sentiment Analysis of Movie Reviews Based on Machine Learning and Deep Learning Models
Shida Yan, Shaochen Cui · 2025
Sentiment analysis has become increasingly important in natural language processing, yet fine-grained sentiment classification remains challenging due to data sparsity and class imbalance. To address these gaps, this study leverages the Movie Reviews Dataset from NLTK, transforming binary sentiment labels into five fine-grained categories—“Very Negative,” “Negative,” “Neutral,” “Positive,” and “Very Positive”—through a custom scoring mechanism. The objective is to evaluate the effectiveness of traditional machine learning models and deep learning models in fine-grained sentiment classification. This research employs five models—Support Vector Machine (SVM), Random Forest, Logistic Regression, LSTM, and CNN—with evaluation metrics such as classification reports, confusion matrices, ROC curves, and AUC scores. The results indicate that SVM and Logistic Regression are well-suited for small-scale, high-dimensional sparse datasets, achieving average AUC scores of 0.75 and 0.72, respectively. LSTM and CNN showed promise for learning intricate textual features but suffered from class imbalance, inhibiting their full potential. The results underscore the robustness of classic ML models on small datasets and the need for larger data and optimized approaches for deep learning models' potential optimizations. The follow-up study discusses data augmentation, sophisticated feature representation techniques, and improved deep learning architectures for advances in sentiment classification tasks.