Sentiment Analysis in Low-Resource Bangla Text Using Active Learning

Md Afnan Ul Haque, Ashiqur Rahman, M. M. A. Hashem · 2021

Sentiment analysis has become a very crucial and extensive dispute in the field of natural language processing and data mining. In supervised learning methods, it needs a huge number of annotated training examples for training. But it is time-consuming, costly, and takes a huge effort to build an annotated dataset. Bangla as a low resource language, it is more adamant to get an annotated dataset for sentiment analysis. In this paper, we proposed an active learning method for sentiment analysis in the Bangla language to reduce the cost of data annotation, preserving classification accuracy. We conducted the feasibility study of active learning with feature extraction methods like TF-IDF and Word Embedding with classifiers like Support Vector Machine(SVM), Logistic Regression(LR), Artificial Neural Network(ANN) along with multiple sample selection strategies on our large dataset. Our result shows active learning can reduce data annotation costs and increase accuracy significantly.

Read the paper · More papers on PaperTik