Combining PhoBERT and SentiWordNet for Vietnamese Sentiment Analysis
Hong-Viet Tran, Van-Tan Bui, Dinh-Tien Do, Van-Vinh Nguyen · 2021
Sentiment analysis is one of the most important NLP tasks, where machine learning models are trained to classify text by polarity of opinion. Many models have been proposed to tackle this task, in which pre-trained PhoBERT models are the state-of-the-art language models for Vietnamese. PhoBERT pre-training approach is based on RoBERTa which optimizes the BERT pre-training method for more robust performance. In this paper, we introduce a new approach to combine phoBERT and SentiWordnet for Sentiment Analysis of Vietnamese reviews. Our proposed sentiment analysis model using PhoBERT for Vietnamese, which is a robust optimization for Vietnamese of the prominent BERT model, and SentiWordNet, a lexical resource explicitly devised for supporting sentiment classification applications. Experimental results on the dataset VLSP 2016 and AIVIVN 2019 demonstrate that our sentiment analysis system has achieved good performance in comparison to other models.