A Vietnamese Answer Extraction Model Based on PhoBERT
Chinh Trong Nguyen, Dang Tuan Nguyen · 2021
PhoBERT pre-trained models have shown its outperformance in many natural language processing tasks. Fine-tuning PhoBERT models is possibly the efficient way to build Vietnamese deep models for answer extraction. For building a Vietnamese answer extraction model using PhoBERT pre-trained model, we need a large SQuAD style annotated dataset. However, there are existing English annotated datasets for answer extraction task and multilingual BERT models which are possibly fine-tuned on English dataset and used in other languages. Therefore, we would find a pre-trained model and a way of fine-tuning this pretrained model for Vietnamese answer extraction task with a low cost of building Vietnamese annotated dataset. We have conducted the experiments with multilingual BERT pre-trained model and PhoBERT pre-trained model to show the performance of these pre-trained models. In the experiments, we have used Vietnamese translated version of SQuAD dataset and Vietnamese manually annotated dataset to show whether the Vietnamese translated dataset is useful in building an answer extraction model. Our experiment results showed that a PhoBERT pre-trained model is a good choice for building a Vietnamese answer extraction model.