Study on Phrase-Window-based Annotation Rules and an Identification Algorithm

Lei Zhang, Li Jian, Haitao Wang, Hantao Tao, Dawei Wu, Liu Yijian · 2021

Word segmentation is heavily used during grammatical dependency analysis in Natural Language Processing. Most of the analysis uses the end-to-end deep learning method based on supervised learning. There are two main issues in this method: firstly, complicated annotation rules result in considerable annotation difficulties and a large workload; secondly, the algorithm cannot identify the multi-granular and diverse features of language. To solve these two issues, we propose phrase-window-based annotation rules and design the Syntax Window Model. Considering phrases as the minimum units, the annotation rules classify sentences into seven categories of nestable phrases and mark the syntactic dependencies between the phrases at the same time. Syntax Window Model, which draws ideas from the Target Detection in Computer Vision, can identify the starting and ending positions of various phrases in sentences, and simultaneously recognize nested phrases and syntactic dependencies. The results show that the annotation rules can be used easily without ambiguity. Compared with the end-to-end algorithm, the algorithm is more suitable for the multi-granularity and diversity in grammar. Experiments on Chinese Phrase Window Dataset (CPWD) using our algorithm improved accuracy by about one percent when compared with the end-to-end method. The model was applied to a CCL2018 competition, and won first place in the analysis of Chinese emotion metaphors.

Read the paper · More papers on PaperTik