Word and Sentence Tokenization with Hidden Markov Models
Bryan Jurish, Kay-Michael Würzner · LDV-Forum/Journal for language technology and computational linguistics · 2013
Word and Sentence Tokenization with Hidden Markov ModelsWe present a novel method ("waste") for the segmentation of text into tokens and sentences.Our approach makes use of a Hidden Markov Model for the detection of segment boundaries.Model parameters can be estimated from pre-segmented text which is widely available in the form of treebanks or aligned multi-lingual corpora.We formally define the waste boundary detection model and evaluate the system's performance on corpora from various languages as well as a small corpus of computer-mediated communication.