Hybrid Method for Sentiment Analysis Using Homogeneous Ensemble Classifier
Murni Murni, Tri Handhika, Achmad Fahrurozi, Ilmiyati Sari, Dewi Putrie Lestari, Revaldo Ilfestra Metzi Zen · 2019
We analyzed the sentiment of a TripAdvisor's review as one of the leading tourist review webs. In TripAdvisor, the sentiment of travelers to the attraction they reviewed is described by a star rating with a numerical scale that ranks from 1 (Terrible) to 5 (Excellent). Travelers usually use this as their polarity. However, there is often inconsistency between their comments and their polarity score such that it cannot be representative of all comments. We used hybrid method to be an alternative solution for this problem which combines two original method, i.e. the lexical-based method and the machine learning-based method. The SenticNet in the lexical-based method is used to determine the representative label of each review. As for the machine learning-based method, we tried to apply several homogeneous ensemble classifiers, such as Bagged Decision Trees, Logistic Model Tree, and Bagged Multi-layer Perceptron algorithms, to determine the sentiment of either existing or new review data with some selected features. The experimental results show that the Homogeneous Ensemble Random Forest algorithm provides the highest accuracy, i.e., 98.1267%. Finally, we obtained the positive review percentages for each attraction in Bali as an alternative for travelers to get information about the sentiment of some attractions other than the unaffected subjective star rating. Some attractions have slightly different between sentiments of hybrid-based methods proposed and star rating, while others have a significant difference.