Research on Chinese Medical Named Entity Recognition Based on ALBERT and IDCNN
Ziyue Zhang, Jin Li, Yan Huang, Weilin Li · Atlantis Highlights in Intelligent Systems/Atlantis highlights in intelligent systems · 2022
BERT (Bidirectional Encoder Representations from Transformers) as a pre-training model has been widely used in the field of natural language processing, of course, it also covers the field of Chinese medical text.In the process of actually dealing with Chinese tasks, BERT also has its own shortcomings, including the lack of Chinese word segmentation.This is because BERT is segmented based on the granularity of words.In addition, the amount of pre-training parameters of the BERT model is too large, which will also cause some problem of poor model performance caused by excessive computing power requirements, long training time, and excessive parameters.To solve the above problems, this paper proposes a Chinese medical named entity recognition model based on ALBERT and IDCNN.Experiments show that the ALBERT-IDCNN-CRF model constructed in this paper has a good performance on the Chinese electronic medical record named entity recognition task, and effectively solves the problems of polysemy and word recognition completion in Chinese electronic medical record named entity recognition.On the CCKS 2017 dataset the model effect F1 value reached 94.51%, and on the CCKS 2019 dataset, the model effect F1 value reached 88.61%.