Text Classification Using Machine Learning Techniques: Comparative Analysis
Ankita Sinha, M. Nazma B. J. Naskar, Manjusha Pandey, Siddharth Swarup Rautaray · 2022
Sharing short reviews on Social Networking Sites has contributed to rising textual data. For this short review text, needed a text classification model to identify short texts, organize them in structured form to define predefined classes efficiently and accurately. In this paper, Imdb movie review text classification model is designed. The Proposed paper is categorized into three phases: - Pre-processing of textual data, Weighting of text, Development of classifier model. For words weighting process, we considered term frequency-inverse document frequency (Tfidf), N-grams (Bi-grams) and term frequency-inverse document frequency (Tf-idf) with N-grams (Bi-grams). After the phases, evaluation of the classification model is done by comparing different algorithms of classification like Naive Bayes, Support vector machine (SVM), Logistic regression, k-nearest neighbours (k-NN) and Decision tree on statistical parameters like accuracy, f1 score, precision, confusion matrix, recall. Finally, the liberation of each weight processing with the classifier is discussed. Combination approach Tf-idf with Bi-grams perform least in every data split scenario and with every classifier. Combinational approach gives negative results as compared to the other two weighting processes. However, the highest accuracy is Logistic Regression with Ngrams (Bi-grams) and Support Vector Machine (SVM) with term frequency-inverse document frequency (Tf-idf) in an 80:20 data split scenario with every classifier.