Chinese named entity recognition based on MacBERT and Joint Learning

Dangguo Shao, Kun Huang, Lei Ma, Sanli Yi · 2023

Extracting medical entities from Chinese medical texts is of great significance to the establishment and application of medical information systems. Chinese text lacks spacers to segment words, which makes it difficult to recognize entities. Therefore, this paper proposes a named entity recognition method combining MacBERT and Joint Learning. The MLM as correction BERT (Mac-BERT) pretrained language model, which alleviates the differences between pre-training and fine-tuning stages, obtains dynamic word vector feature expression, and then enters into a framework that integrates the named entity recognition task and the word segmentation task for joint learning, so as to improve the feature capture ability of the model. Experiments show that the model can effectively obtain the important entities in the instruction of traditional Chinese medicine. On the Chinese medicine instruction dataset, fusing MacBERT with Joint Learning model achieved the best result with 68.43% F1, which increased by 1.6% compared with the two-layer BiLSTM non-joint model, and increased by 1.07% compared with the combined BERT with Joint Learning model. Experimental results show that the recognition performance is better than that of the non-joint model and other pretrained models.

Read the paper · More papers on PaperTik