Named Entity Recognition in Chinese Medical Texts Based on RoBERTa-WWM-IDCNN-CRF
Lan Yu Tang, Jie Kong, Liang Xu · 2024
Electronic medical records and doctor-patient conversations contain a wealth of useful information, such as disease symptoms, drug names, and cure cycles. Traditional deep learning approaches utilize bidirectional recurrent neural networks to encode text, which cannot effectively capture long-distance dependencies in text. In order to identify medical entity types in various texts more accurately, a named entity recognition approach based on RoBERTa-WWM-IDCNN-CRF is proposed. This approach introduces the RoBERTa-WWM pre-trained model, in order to improve the overall understanding of long Chinese texts for the first time. The model generates semantic representations containing prior knowledge. Then, the generated semantic representations are feed to the IDCNN to extract sentence level characteristics. Finally, the global optimal solution is obtained and the results of the named entity are produced through the CRF. The experiment indicates that the improvement of $F_{1}$ score, accuracy, and recall of RoBERTa-WWM-IDCNN-CRF model are 7.94%, 4.52% and 10.02% comparing with the classical BiLSTM-CRF model.