Comparison of Machine Learning Text Classification for Intent Sentiment Analysis
Abulwafa Muhammad, Sarjon Defit, Gunadi Widi Nur Cahyo · 2023
Tourist destination reviews on Google Maps have become a valuable point of reference for visitors seeking enjoyable spots to visit. Additionally, users can gain insight into the reasons for writing reviews. Text classification is used to determine the motivation behind these reviews. The aim of this study is to compare machine learning (ML) text classification techniques for identifying the intent behind sentiments expressed as complaints (0), suggestions (1), opinions (2), statements(3), and awards (4). The ML techniques assessed are Multinomial Naive Bayes (MNB), Support Vector Machine (SVM), and K-Nearest Neighbor (KNN). To achieve this, tourist destination review data sets, comprising 738 reviews, were collected using web scraping from the Google Maps website. Preprocessing stages, including case-folding, tokenisation, filtering and stemming, are executed prior to the data being prepped for feature extraction with TF-IDF. The employed classification models perform well with KNN achieving 1.00 accuracy on a 90:10 data split, and SVM obtaining 0.90 on the same data split. Although, the Multinomial Naive Bayes algorithm with 0.70 accuracy on a 90:10 data split is classified as fair. Comparison of methods was conducted by means of a Receiver Operating Characteristic (ROC) diagram. The testing accuracy for the SVM method was found to be 95%, MNB at 91%, and KNN at 84%.