PEM: A Medical Named Entity Recognition Method Based on Proximity Enhancement
Mei Liu, Hongyun Huang, Zuohua Ding · 2024
Named entity recognition is the most basic task in natural language processing, and its quality directly affects the performance of downstream tasks. Due to the scarcity of annotated corpus and diverse species in the medical domain, as well as the continuous emergence of neologisms, the application of general domain entity recognition methods in vertical fields face the problem of inaccurate recognition of professional vocabulary. Proximity relation as a strong prior information is of great significance for the annotation task of named entity recognition. In view of the insufficient of proximity modeling in current work, we propose a named entity recognition model based on proximity relationship enhancement (PEM). Firstly, BERT is introduced to obtain the semantic features of text sequences. At the same time, the embedding vectors corresponding to lexical labels and co-occurrence matrices are input into the graph attention network. This network uses an information gate mechanism to model the proximity dependencies and filter the redundant features. Furtherly, we utilize the multi-head cross attention mechanisms to align and fuse the proximity features, and introduce BiLSTM to extract the contextual semantics. Finally, the fused features are fed into CRF for decoding. In order to verify the effectiveness of the model in this paper, comparative ablation experiments are carried out on four commonly used medical entity datasets, such as CCKS2018. The experimental results further demonstrate that the PEM model outperforms other recognition models and can effectively improve the accuracy of entity recognition with good robustness.