Construction of Chinese Social Media Named Entity Dataset with Multi-Tool Fusion Annotation
Hechuang Wang, Shixuan Peng, Guojun Huang, Yu Wang, Yaqiong Qiao · 2024
Information extraction from social media platforms is crucial for the advancement of downstream natural language processing tasks. However, the limitation of NER data in Chinese social media significantly impedes the progress of NER tasks in the Chinese domain. To tackle the scarcity of fine-grained chinese named entity datasets in social media and the annotation challenge, we propose an automatic annotation method based on three available named entity recognition tools, and construct a new Chinese social media named entity dataset (named WBNER) using this method. The automated annotation method comprises two main components. Initially, we construct a fine-grained dictionary to validate the effectiveness of entity extraction by named entity recognition tools, which focuses on entity types relevant to social media users and filtered through big data retrieval techniques; Subsequently, we build a corpus entity dictionary using entity recognition tools on the social media corpus, which is then merged with the fine-grained dictionary to form a comprehensive annotated set of corpora. To assess the effectiveness of our method, we construct a 10-entity-type fine-grained named entity dataset for Weibo and compare it with existing chinese named entity datasets. Experimental evaluations on baseline models for sequence tagging show that the dataset constructed using this method can effectively perform relevant tasks. Therefore, the automated named entity recognition annotation method reduces manual workload while maintaining accuracy, and we will release WBNER at https://github.com/Superpsx/WBNER.