Vietnamese news classification based on BoW with keywords extraction and neural network

Toàn Phạm Văn, Tạ Minh Thanh · 2017

Nowadays, text classification (TC) becomes the main applications of NLP (natural language processing). Actually, we have a lot of researches in classifying text documents, such as Random Forest, Support Vector Machines and Naive Bayes. However, most of them are applied for English documents. Therefore, the text classification researches on Vietnamese still are limited. By using a Vietnamese news corpus, we propose some methods to solve Vietnamese news classification problems. By employing the Bag of Words (BoW) with keywords extraction and Neural Network approaches, we trained a machine learning model that could achieve an average of ≈ 99.75% accuracy. We also analyzed the merit and demerit of each method in order to find out the best one to solve the text classification in Vietnamese news.

Read the paper · More papers on PaperTik