Sentence-Level Resampling for Named Entity Recognition
Xiaochen Wang, Yue Wang · Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies · 2022
As a fundamental task in natural language processing, named entity recognition (NER) aims to locate and classify named entities in unstructured text.However, named entities are always the minority among all tokens in the text.This data imbalance problem presents a challenge to machine learning models as their learning objective is usually dominated by the majority of non-entity tokens.To alleviate data imbalance, we propose a set of sentence-level resampling methods where the importance of each training sentence is computed based on its tokens and entities.We study the generalizability of these resampling methods on a wide variety of NER models (CRF, Bi-LSTM, and BERT) across corpora from diverse domains (general, social, and medical texts).Extensive experiments show that the proposed methods improve performance of the evaluated NER models especially on small corpora, frequently outperforming sub-sentence-level resampling, data augmentation, and special loss functions such as focal and Dice loss. 1 Domain NER Corpus # of Tokens # of Entity Types % of Least vs.Most Freq.Type Entity Tokens % of All Entity Tokens # of Sent.% of Sent.w/ Entities Social WNUT 59,570 6 0.43 vs. 1.59 5.03 3,394 36.18General GMB subset 66,161 8 0.03 vs. 4.00 15.03 2,999 85.40 Medical AnEM 71,697 11 0.03 vs. 1.08 3.91 2,815 35.38 Medical CADEC 121,307 5 0.21 vs. 6.65 15.76 5,719 58.86 General CoNLL 204,567 4 2.26 vs. 5.46 16.64 14,986 74.28 Medical n2c2 ADE 813,277 9 0.19 vs. 2.34 10.89 65,293 22.73 General OntoNotes 2,200,865 18 0.01 vs. 2.59 10.89 115,812 50.11