Pause and Stop Labeling for Chinese Sentence Boundary Detection

Hen‐Hsen Huang, Hsin‐Hsi Chen · 2011

The fuzziness of Chinese sentence boundary makes discourse analysis more challenging. Moreover, many articles posted on the Internet are even lack of punctuation marks. In this pa-per, we collect documents written by masters as a reference corpus and propose a model to label the punctuation marks for the given text. Conditional random field (CRF) models trained with the corpus determine the correct delimiter (a comma or a full-stop) between each pair of successive clauses. Different tag-ging schemes and various features from differ-ent linguistic levels are explored. The results show that our segmenter achieves an accuracy of 77.48 % for plain text, which is close to the human performance 81.18%. For the rich for-matted text, our segmenter achieves an even better accuracy of 82.93%. 1

Read the paper · More papers on PaperTik