Employing Machine Learning techniques on Sentiment Analysis of Google Play Store Bangla Reviews

Md Muhtasim Jawad Soumik, Syed Salvi Md Farhavi, Farzana Eva, Tonmoy Sinha, Mohammad Shafiul Alam · 2019

This article offers an in-depth insight on a number of existing methodologies to perform sentiment analysis using text classification on Bangla dataset. Although the rapidly developing machine learning algorithms are showing promising results, the viability of those methods for non-English languages such as Bangla is yet to be fully explored. This research aims to fill in some of those existing research gaps through proper implementation of machine learning techniques where words are converted into feature vectors via implementation of TF-IDF algorithm on data crawled from Google play store, the largest Android application market. Many significant algorithms staring linear algorithms like Naïve Bayes, Linear Support Vector Machine (SVM) are implemented. An in-depth comparison is also made among the results of various existing algorithms. The experimental results indicate that even the base-line algorithms, after proper pre-processing, can show promising results on our Bangla dataset. Naïve Bayes, Support Vector Machine and Logistic Regression has shown very promising results (accuracy score of 0.75 on average) even with the data limitation. An Ensemble method is also proposed with Adaptive Boosting technique showing an accuracy score of 0.7639 with five-fold applied. SVM has the best accuracy score of 0.7648 among all the algorithms when five-fold is applied and Gradient Boosting has the best accuracy score of 0.7695 when five-fold is not applied.

Read the paper · More papers on PaperTik