Discriminative Learning with Natural Annotations: Word Segmentation as a Case Study

Wenbin Jiang, Meng Sun, Yajuan Lü, Yating Yang, Qun Liu · 2013

Structural information in web text provides natural annotations for NLP problems such as word segmentation and parsing. In this paper we propose a discriminative learning algorithm to take advantage of the linguistic knowledge in large amounts of natural annotations on the Internet. It utilizes the Internet as an external corpus with massive (although slight and sparse) natural annotations, and enables a classifier to evolve on the large-scaled and real-time updated web text. With Chinese word segmentation as a case study, experiments show that the segmenter enhanced with the Chinese wikipedia achieves significant improvement on a series of testing sets from different domains, even with a single classifier and local features. 1

Read the paper · More papers on PaperTik