Research on Entity Relation Extraction of Chinese Medical Texts Based on Pre-Training Model
Shuang Chen, Qun Hou, Ying Chen · 2024
Aiming at the problems of intensive knowledge in Chinese medical field, more specialized vocabulary and limited manual annotation data set, a framework for entity relation extraction of Chinese medical texts based on pre-training model was proposed. Firstly, Masking language modeling (MLM) was taken as a pretraining task, and the robust optimized bidirectional encoder model (Roberta) was continued to be pre-trained using Chinese medical corpus, and the weight of the pre-trained model in the field of Chinese medical texts was adjusted to improve the pre-trained model's understanding of Chinese medical vocabulary. Secondly, based on the adjusted pretraining model, the entity relation extraction task is added, and the model is fine-tuned to obtain the entity relation extraction model of Chinese medical text. In the Chinese medical text entity relation extraction dataset CMeIE-V2, the entity relation extraction accuracy of this model reached 60.38% and F1 value reached 54.49%. Four pre-trained models, BERT-large, BioBERT, MCBERT and RoBERTa-large, were compared. Compared to the best-performing RoBERTa model, the F1 value is improved by 0.8% and the accuracy is improved by 0.9%. The experimental results show that the pre-training model trained with Chinese medical field corpus can effectively improve the effect of entity relation extraction in Chinese medical field.