Detecting Segmentation Errors in Chinese Annotated Corpus

Chengjie Sun, Changning Huang, Xiaolong Wang, Mu Li · 2005

This paper proposes a semi-automatic method to detect segmentation errors in a manually annotated Chinese corpus in order to improve its quality further. A particular Chinese character string occurring more than once in a corpus may be assigned different segmentations during a segmentation process. Based on these differences our approach outputs the segmentation error candidates found in a segmented corpus and then on which the segmentation errors are identified manually. Segmentation error rate of a gold standard corpus can be given using our method. In Peking

Read the paper · More papers on PaperTik