LACNER: Enhancing Few-Shot Named Entity Recognition with Label Words and Contrastive Learning
Yuhui Xiao, Qun Yang, Jianjian Zou, Sichi Zhou · 2024
Few-shot named entity recognition aims to extract entities by using a limited amount of annotated data. Recently, contrastive learning shows promising performance in few-shot named entity recognition. Despite the effectiveness of this method, there are still a few obstacles for few-shot named entity recognition. Firstly, contrastive learning mainly focuses on optimizing the discrimination of tokens, which offers limited optimization for encoders, causing suboptimal entity representations. Secondly, the scarcity of data may lead to model overfitting during the training process. To resolve these issues, we propose a contrastive learning method combined with label words, providing extra information to our model in order to optimize token representations. At same time, we use dropout noise as data augmentation, generating sentences semantically similar to the original ones. These sentences increase the diversity of samples. Moreover, due to the inconsistency of dropout in the training and prediction stages, it thereby harms the generalization of the model. Therefore, we alleviate it by using bidirectional Kullback–Leibler divergence, which constrains the output of model under different dropouts to be essentially consistent. Our extensive experiments on multiple datasets show that LACNER achieves new state-of-the-art results in most cases.