Corpus Annotation System Based on HanLP Chinese Word Segmentation

Xuanjun Liu, Zheyu Zhu, Tengyan Fu, Jiaxuan Chen, Ying Jiang · The 2nd International Conference on Computing and Data Science · 2021

The Guangdong-Hong Kong-Macao Greater Bay Area is one of the most open and economically dynamic areas in China. It has been playing an important strategic role in the overall national development. The diversified languages and cultures in Hong Kong, Macau and Guangdong makes language problems much more complicated. Natural language processing based on deep learning has become an important method to construct and develop standardized semantic rules for language resources in Guangdong, Hong Kong and Macao. Therefore, a large amount of correct corpus data has become the basis for solving such problems. The corpus annotation system based on HanLP Chinese word segmentation performs preliminary annotation on the uploaded corpus, which is further proofread by annotators. In addition, the system cross-validate and submit the annotation results to output high-quality corpus. The corpus annotation system consists of a user module and an administrator module, with functions such as full-text corpus retrieval and data statistics based on Elasticsearch. As a natural language processing tool, HanLP Chinese word segmentation is used for storing the annotated corpus word segmentation in the system, which has certain advantages in corpus analysis and mining. This corpus annotation system can provide a large amount of high-quality corpus data for the semantic analysis of the Guangdong-Hong Kong-Macao Greater Bay Area.

Read the paper · More papers on PaperTik