Sentiment Analysis Using Part-of-Speech-Based Feature Extraction and Game-Theoretic Rough Sets
Yixing Chen, JingTao Yao · 2021 International Conference on Data Mining Workshops (ICDMW) · 2021
Sentiment analysis, one of the most trending natural language processing tasks, is used to mine opinions or sentiments from a given text. Two significant challenges of sentiment analysis are 1) complexity in data pre-processing caused by the high dimensionality of textual data; 2) uncertainty in classifying sentiment polarities due to the ambiguity of natural languages. To address these issues, we propose a model using part-of-speech-based feature extraction to reduce dimensionality and game-theoretic rough sets (GTRS) to establish a balance between the accuracy and coverage trade-off. We evaluate this model with three different sizes of datasets (Yelp reviews, IMDB movie reviews, and Amazon product reviews). The experiment results show that the proposed model outperforms Pawlak’s rough set model and 0.5-probabilistic rough set model. In comparison with four traditional binary classification models (i.e., SVM, naïve Bayes, decision tree, and KNN), the proposed model also achieves higher accuracy rates. This research suggests that the proposed model is promising to deal with the complexity and uncertainty in sentiment analysis tasks.