Improved Pre-training and Semi-supervised Learning for Domain-specific Chinese Named Entity Recognition

Bin Wang, Dafu Tang, Qing Wu · 2022

With the explosive growth of information, automated text mining and knowledge discovery have become a hot spot in the research of Artificial Intelligence (AI). Named Entity Recognition (NER) is an important AI technology that needs to extract important entity information from the corresponding text. Today, most of the current NER technologies are based on Chinese characters. These methods, however, greatly ignores the vocabulary information in Chinese. How to use the rich vocabulary information in Chinese to enhance the effect of named entity recognition in the specific domain is still a challenge. What's more, in different domain, the labeled corpus is scarce, and labeling is expensive. Based on the pre-training model BERT and a large number of unsupervised corpus that are easy to obtain in the domain, this paper proposes an improved pre-training method of whole word mask of unsupervised corpus in the domain, and fully uses the unsupervised corpus in the fine-tuning task stage for semi-supervised learning. The experiment is carried out on the datasets in the domains of law and traditional Chinese medicine. The results show that the performance of the model method proposed in this paper in the domain dataset greatly exceeds that of the baseline model.

Read the paper · More papers on PaperTik