An Ensemble Method for Sentiment Classification of Long Vietnamese Documents

Ngoc Toan Nguyen, Xuan Tuan Le, The Dung Luong · 2022 RIVF International Conference on Computing and Communication Technologies (RIVF) · 2022

Sentiment classification of text is one of the major tasks of natural language processing which has has gained significant interest recently. While there are many researches in this field dealing with many other languages, it is quite chal-lenging to handle Vietnamese text, especially in case of specific domains such as social security protection domain where the classification requires and depends much on expertise knowledge of that field to collect and to label the dataset. In this work, we collect 2541 Vietnamese news (a paragraph or a whole text document) from websites and social networks such as Facebook to build our dataset and propose an ensemble model of Support Vector Machine (SVM) and Light Gradient Boosting Machine (LightGBM) to classify whether the news is normal, negative or anti-social. To evaluate the performance of our propose model, we also experiment on the same dataset with other state of the art classification models such as Naive Bayes, Maximum Entropy, PhoBERT, XLM-RoBERTa. Preliminary results show that our propose model outperforms those model in term of accuracy and f1 score.

Read the paper · More papers on PaperTik