Prompt-Based Data Augmentation Using Contrastive Learning Under Scarcity of Annotated Data
Muhammad Uzair Ul Haq, Davide Rigoni, Alessandro Sperduti · Frontiers in artificial intelligence and applications · 2024
Named Entity Recognition is a crucial task in Natural Language Processing (NLP) which aims to identify the entities in text. Given an adequate amount of annotated data, Large Language Models (LLMs) have been shown to be effective in this task when fine-tuned. However, the performance of LLMs is severely affected when annotated datasets are limited. To alleviate this problem, adding synthetic data via Data Augmentation (DA) techniques is a viable approach. Even so, DA for token-level tasks suffers from two main limitations: (i) token-label misalignment problem; and (ii) quality of generated synthetic data. In this paper, we propose a novel prompt-based DA approach using contrastive learning. The proposed method can generate high-quality synthetic data while preserving the token-label correspondences. Experimental results demonstrate that the proposed approach, when compared against multiple baselines on well-known Named Entity Recognition (NER) datasets, achieves State-of-the-Art performance.