Chinese Open-domain Named Entity Boundary Identification based on A Self-Training Method

FU Ruij · Intelligent Computer and Applications · 2014

Named entity recognition is an important task in the domain of Natural Language Processing,which plays an important role in many applications. This paper focuses on the boundary identification of Chinese open-domain named entities. Because the shortage of training data and the huge cost of manual annotation,the paper proposes a self-training approach to identify the boundaries of Chinese open-domain named entities in context. Due to the lack of training data,the paper firstly generates a large scale Chinese proper noun corpus based on parallel corpora,and also transforms a Chinese dependency tree bank to a noun compound training corpus. Subsequently,the paper proposes a self-training-based approach to combine the two corpora and train a model to identify boundaries of named entities. The experiments show the proposed method can take full advantage of the two corpora and improve the performance of named entity boundary identification.

Read the paper · More papers on PaperTik