A Study on Named Entity Recognition of Chinese Social Text Based on ERNIE and Lexicon Enhancement

Xiuli Li, Yi Wang, Yaqiong Qiao, Yu Wang, Jiahui Li · 2024

Named Entity Recognition (NER) is a task in natural language processing that focuses on identifying and categorizing significant entities within a body of text. The efficiency of this process has a ripple effect on subsequent operations, including tasks like extracting information, answering questions, and translating languages, among others. When pitted against English's approach to named entity recognition, A significant hurdle in the recognition of Chinese named entities is the unpredictability inherent in Chinese linguistic structures, coupled with the diverse and mutable ways in which words are strung together to form phrases, especially in more colloquial datasets, such as social texts. Dealing with these complex datasets requires more accurate algorithmic models, and semantic elements, including the broader context, must be factored in. As a result, the article introduces an advanced Chinese named entity recognition model known as LE-ERNIE, which incorporates lexicon enrichment.Lexicon enhancement, on the other hand, aims to incorporate information that cannot be directly mined from the text through an external dictionary. By employing an adaptive coding mechanism, the model merges lexicon details into its encoding phase, facilitating the assimilation of profound lexical insights within the pre-established framework of ERNIE's language processing capabilities. The final comparative experiments, based on the Weibo NER dataset, showed the following results: ① Compared to the ERNIE model without lexicon enhancement, the proposed LE-ERNIE model achieved an improvement of approximately 0.95% in F1-score. ② Compared to other models with lexicon enhancement, the LE-ERNIE model demonstrated an F1-score improvement ranging from 0.71% to 12.67%.

Read the paper · More papers on PaperTik