Vietnamese news classification based on BoW with keywords extraction and neural network
Toàn Phạm Văn, Tạ Minh Thanh · 2017
Nowadays, text classification (TC) becomes the main applications of NLP (natural language processing). Actually, we have a lot of researches in classifying text documents, such as Random Forest, Support Vector Machines and Naive Bayes. However, most of them are applied for English documents. Therefore, the text classification researches on Vietnamese still are limited. By using a Vietnamese news corpus, we propose some methods to solve Vietnamese news classification problems. By employing the Bag of Words (BoW) with keywords extraction and Neural Network approaches, we trained a machine learning model that could achieve an average of ≈ 99.75% accuracy. We also analyzed the merit and demerit of each method in order to find out the best one to solve the text classification in Vietnamese news.