Automatic Identification of Chinese Multiword Chunk Based on CRF
Ru Li, Lijun Zhong, Shuanghong Li, Zezheng Zhang · 2010
Identifying the Chinese multiword chunk automatically is a newly emerged technology in the NLP field. As anew strategy, it can effectively improve the performance of the syntactic parsing. The work follows the standard description system of Chinese multiword chunk and has constructed two tag sequence models based on CRF model, which are named as ”the syntactic mark tagging list model” and ”the sequence mark tagging list model” respectively. The corpus used in the training process is called as ”the Chinese multiword chunk bank”, which is provided by Tsinghua University. In the experiments, by selecting appropriate the features and introducing some important rules, the better results are achieved and this system for identifying the Chinese multiword chunk can run well in a restricted area. Thus, it provides a bridge between syntax and semantic content.