Automatic Bengali Document Categorization Based on Word Embedding and Statistical Learning Approaches

Md. Rajib Hossain, Mohammed Moshiul Hoque · 2018 International Conference on Computer, Communication, Chemical, Material and Electronic Engineering (IC4ME2) · 2018

The automated categorization of text documents into predetermined categories has witnessed a growing in the last few years, due to the huge availability of documents in digital form and the ensuing need to organize them. Automatic document categorization is the process of assigning one or more categories or classes to a document, making it easier to manipulate and sort. This paper proposes a Bengali document categorization technique based on word2vec word embedding model and stochastic gradient descent (SGD) statistical learning algorithm with multi-class svm. The semantic features of a document are extracting by Word2Vec and SGD improve the classification complexity with multi-class SVM that classify the unlabeled data. The experimental result with 10000 training and 4651 testing documents shows the 93.33% accuracy.

Read the paper · More papers on PaperTik