Bangla Document Categorisation using Multilayer Dense Neural Network with TF-IDF
Manisha Chakraborty, Mohammad Nurul Huda · 2019 1st International Conference on Advances in Science, Engineering and Robotics Technology (ICASERT) · 2019
Document categorization is a classic NLP task which includes allocating documents based on their content into one or more predefined classes. In this paper, we proposed a model using multilayer dense neural network with TF-IDF as feature selection technique in terms of Bangla text document categorization. The proposed system is distributed into three consecutive steps: i) preprocessing raw text data and extracting feature using TF-IDF, ii) designing the model architecture and fitting the model to train set, and iii) evaluating model performance on test set based on accuracy and F1 score. It is observed from experiments that the proposed method exhibits higher accuracy (84.58%) and F1 score (0.84) compared to the other well-known classification algorithms (Support Vector Machine, Multinomial Naive Bayes and so on) for Bangla text document classification.