Unsupervised Learning of Word Boundary with Description Length Gain

Chunyu Kit, Yorick Alexander Wilks · 1999

This paper presents an unsupervised approach to lexical acquisition with the goodness measure description length gain (DLG) formulated following classic information theory within the minimum description length (MDL) paradigm. The learning algorithm seeks for an optimal segmentation of an utterance that maximises the description length gain from the individual segments. The resultant segments show a nice correspondence to lexical items (in particular, words) in a natural language like English. Learning experiments on large-scale corpora (e.g., the Brown corpus) have shown the effectiveness of both the learning algorithm and the goodness measure that guides that learning. 1. Introduction Detecting and handling unknown words properly has become a crucial issue in today's practical natural language processing (NLP) technology. No matter how large the dictionary that is used in a NLP system, there can be many new words in running/real texts, e.g., in scientific articles, newspapers and We...

Read the paper · More papers on PaperTik