Hotel Review Sentiment Analysis Using Indonesian Language Based on Machine Learning

Nurhafnita, Ramzi Adriman, Taufik Fuadi Abidin · 2022

It is crucial to optimize the Naive Bayes technique because its level of accuracy still has flaws. In order to achieve a higher level of accuracy, optimization employs the right and best techniques for text grouping, particularly for hotel review classification. In order to increase the precision of sentiment analysis, this study compares the use of a dataset with 6 features and a dataset with 18 features in order to determine the impact on the classification accuracy of Naive Bayes with Chi-square. This is demonstrated when the Naive Bayesian algorithm is applied to a dataset with 18 features, which results in accuracy values of f-measure = 0.83 and ROC = 0.922 as opposed to 6 features, which have f-measure = 0.746 and ROC = 0.839. The Chi Square feature selection technique has the effect of improving the Nave Bayes algorithm's classification accuracy of Indonesian hotel review texts. Based on the f-measure and ROC values, which are f-measure = 0.831 and ROC = 0.92 for a dataset with 18 features, the accuracy value will be high if the Naive Bayes algorithm is combined with Chi Square feature selection. According to the results of this study, the accuracy value can be increased in both the 6-feature dataset and the 18-feature dataset by using the Naive Bayes algorithm in conjunction with Chi Square.

Read the paper · More papers on PaperTik