Towards Robust Named Entity Recognition in Bangla With LLMs Based Data Augmentation
Qiong Zeng, Yuanyu Li, Jian Liu · 2024
Named Entity Recognition (NER) is a pivotal task in Natural Language Processing (NLP), focused on identifying specific entities in unstructured text. While significant advancements have been made, most existing methods are primarily designed for resource-rich languages like English, with limited solutions available for resource-poor languages such as Bangla. To address this gap, we propose a large language model (LLM)-based data augmentation technique aimed at enhancing NER for Bangla. Our approach targets two critical aspects of data augmentation, effectively boosting NER performance in low-resource settings. Experimental results from the Bangla Complex Named Entity Recognition Challenge demonstrate the efficacy of our method, which shows a 5% improvement in F1 score over existing techniques. Moreover, these results underscore the potential of our approach as a promising solution for NER in resource-scarce environments.