Research on Shellfish Text Entity Extraction Method Based on Pre-trained Models
Mingtian Yu, Zhenghua Zeng, Hao Tang, Yonghui Zhang, Yuxin Ao, Uzair Aslam Bhatti, Asmaa Fahim · 2024
Shellfish text information has the problems of fuzzy entity boundaries and strong contextual semantic dependency, and a pre-trained model-based entity extraction method for shellfish text is proposed. The method adopts RoBERTa-wwm pre-trained language model for unsupervised pre-training, obtains word vectors with semantic information and inputs them into BiLSTM layer to further capture the semantic features of shellfish text; finally, the optimal label prediction is obtained by using CRF computation. By collecting shellfish-related raw corpus, preprocessing data and text annotation, a corpus for shellfish entity recognition is constructed, and comparison experiments are carried out based on this data. The experiments show that compared with BiLSTM and BiLSTM-CRF, etc., the proposed method can get the best comprehensive recognition effect, and the comprehensive F1 value is as high as 94.81%.