Large Language Model-Driven Semi-Automated Construction of Named Entity Datasets in Specific Knowledge Domain
Yahan Liang, Hua Cao · 2024
Named Entity Recognition (NER) plays a foundational role in Natural Language Processing (NLP), essential for applications such as question-answering systems and text summarization. Within specific domains, high-quality datasets significantly impact the performance enhancement of NER based on pre-trained large language models (LLMs). Traditionally, the development of high-quality NER datasets has been a manual, time-intensive, and expensive process. However, leveraging emerging large language model (LLM) technologies can notably decrease manpower expenses. This paper proposes a semi-automated method that combines the computational capabilities and knowledge bases of LLMs, focusing on constructing NER datasets for specific domains. Initially, the study assesses the role of prompts in LLM for Chinese NER and establishes effective prompt paradigms. Then, by adopting semi-automatic mapping correction techniques, it significantly improves annotation accuracy while substantially reducing manpower costs. Ultimately, a series of validation experiments demonstrate the effectiveness of the proposed method.