Biomedical Named Entity Recognition Based on MCBERT
Sai Wang, Hankiz Yilahun, Askar Hamdulla · 2022
In the biomedical field, the named entity recognition method using static word vectors to represent semantics has low accuracy and cannot completely represent the upper and lower semantic relationship of medical texts. Therefore, a model for biomedical named entity recognition is proposed, which combines the pre training language model MCBERT with CNN, BILSTM and CRF. First, MCBERT is used for semantic extraction to generate dynamic word vectors, then the word vectors are sent to CNN to extract local sequence features, and then LSTM is allowed to extract long-distance sequence features based on local sequence features. Finally, CRF is used to learn the pre and post dependencies of sentences, so as to add some constraints to ensure the correctness of the final prediction results. The average F1 value of the model on the ccks2019 evaluation task “medical entity recognition and attribute extraction data set for Chinese electronic medical records” has reached 82.8%. The experimental results show that the proposed model effectively improves the accuracy of biomedical named entity recognition.