Elephant: Sequence Labeling for Word and Sentence Segmentation

Kilian Evang, Valerio Basile, Grzegorz Chrupała, Johan Bos · 2013

Tokenization is widely regarded as a solved problem due to the high accuracy that rulebased tokenizers achieve.But rule-based tokenizers are hard to maintain and their rules language specific.We show that highaccuracy word and sentence segmentation can be achieved by using supervised sequence labeling on the character level combined with unsupervised feature learning.We evaluated our method on three languages and obtained error rates of 0.27 ‰ (English), 0.35 ‰ (Dutch) and 0.76 ‰ (Italian) for our best models.

Read the paper · More papers on PaperTik