Character Language Models for Chinese Word Segmentation and Named Entity Recognition
Bob Carpenter · 2006
(Alias-i 2006) to Chinese word segmentation and named en-tity recognition. We provide results for the third SIGHAN Chinese language processing bakeoff (Levow 2006). F1 mea-sures on the best performing corpora were.972 for word seg-mentation and.855 for person/location/organization named-entity recognition. 1 Word Segmentation Chinese is written without spaces between words. For the word segmentation task, four training corpora were provided with one sentence per line and a single space character between words. Test data consisted of Chinese text, one sentence per line, without spaces between words. The task is to insert single space characters between the words. For this task and named entity recognition, we used the UTF8-encoded Unicode versions of the corpora converted from their native formats by the bakeoff organizers.