Data Augmentation Based on Word Importance and Deep Back-Translation for Low-Resource Biomedical Named Entity Recognition

Yu Wang, Meijing Li, Runqing Huang · 2024

To tackle the challenges posed by limited annotation resources and the time-consuming and expensive nature of manual annotation in biomedical named entity recognition tasks, this paper proposes a data enhancement method in low-resource biomedical named entity recognition scenarios. This method is based on based on word importance synonym replacement and screening mechanism for deep back-translation. Firstly, the TF-IDF scores are utilized to assess the significance of words, and synonym replacement is conducted for words considered less significant. Next, we introduce a deep back-translation method with a screening mechanism to further improve the data for content with low semantic similarity and poor readability, which is aimed at addressing the issue of increasing syntactic error rate and semantic deviation that arises from relying solely on synonym replacement. The proposed method has been validated using two biomedical named entity recognition datasets, BioNLP13CG and NCBI-disease. The experimental results demonstrate that the data enhancement method proposed in this study significantly enhances the accuracy and F1 scores of the biomedical named entity recognition task in low-resource scenarios.

Read the paper · More papers on PaperTik