Topics in unsupervised language learning

John Goldsmith, Yu Hen Hu · 2007

Language learning is one of the most complex and challenging problems in Artificial Intelligence. Natural languages have several distinct components, each of which presents a difficult learning challenge. These include phonetics and phonology, morphology, syntax, semantics, pragmatics, and discourse. The focus of this dissertation is natural language morphology, which is the study of the internal structure of words, and in particular with methods for automatically learning morphological structure. Much as sentences are composed of structured sequences of words, words are composed of structured sequences of morphemes. In some languages, the morphological structure is complex; in others, it is relatively simple. Based on how much explicit human analysis is integrated into the learning procedure, language learning can be categorized into supervised and unsupervised. Unsupervised learning is the focus in this thesis, which means the language parameters in our models are learned with little or no active human participation, but is instead induced by a system which is language-independent. Whatever is different about two languages must be inferred by the learning algorithm. The principal problems addressed in this dissertation are the following: automatic identification of morphemes of a language, on the basis of a simple text; finding the best analysis of words into morphemes in a language-independent way; discovery of automatically related forms of a morpheme (known as the problem of allomorphy), and how the task of machine translation from one language to another can be improved by a knowledge of the morphologies of source and target language, and how machine translation can in turn improve the analysis of morphology. Much of the work in this dissertation uses Minimum Description Length analysis, which involves finding the least complex way to describe a finite state automaton that generates the observed data and assigns it a relatively high probability.

Read the paper · More papers on PaperTik