Unsupervised Learning of Lexical Information for Language Processing Systems
Graham Neubig · 2012
Natural language processing systems such as speech recognition and ma-chine translation conventionally treat words as their fundamental unit of processing. However, in many cases the definition of a “word ” is not obvious, such as in languages without explicit white space delimiters, in agglutinative languages, or in streams of continuous speech. This thesis attempts to answer the question of which lexical units should be used for these applications by acquiring them through unsupervised learn-ing. This has the potential to lead to improvements in accuracy, as it can choose lexical units flexibly, using longer units when justified by the data, or falling back to shorter units when faced with data sparsity. In addition, this approach allows us to re-examine our assumptions of what units we should be using to recognize speech or translate text, which will provide insights to the designers of supervised systems. Furthermore, as the methods require no annotated data, they have the potential to remove the annotation bot-