An Evaluation of the UIT-VSFC Dataset Using Modern Machine Learning Techniques and Word Embeddings

Quynh Duong Vu Xuan, Kanjana Laosen, Nasith Laosen · 2021

Feedback from students during the course and upon course completion has become a powerful resource to improve teaching quality and enhance student's learning experience. However, the available data for free use is limited, especially for a low-resource language like Vietnamese. Currently, there is only one dataset in the education domain, called the Vietnamese Students’ Feedback Corpus (UIT-VSFC), that has been published for free use. This study therefore aims at evaluating the available corpus to use as a benchmarking dataset for conducting future researches as well as developing real-world applications. In this paper, deep neural network (DNN) and recurrent neural network (RNN) models are developed employing a word embedding method for two different tasks, i.e., topic and polarity classification. The experimental results show that DNN models outperform RNN models with 85.22% (>84.30%) and 88.56% (>86.32%) of accuracy for topic and polarity classification, respectively. Error analysis is conducted to explore the confusion of labeling and imbalance of data in the dataset. Workarounds for solving the problems are presented together with their results.

Read the paper · More papers on PaperTik