Study on n-gram language models for topic and out-of-vocabulary words

Welly Naptali · Toyohashi University of Technology Academic Institutional Repository (Toyohashi University) · 2011

Language models (LMs) are an important field of study in automatic speech recognition (ASR) systems. LM helps acoustic models find the corresponding word sequence of a given speech signal. Without it, ASR systems would not understand the language and it would be hard to find the correct word sequence. A data sparseness problem for modeling a language often occurs in LMs. The problem is caused by the insufficiency of training data, which in turn, makes the infrequent words have unreliable probability. In this research, we investigated a class LM based on a latent semantic analysis (LSA). A word-document matrix is commonly used to represent a collection of text (corpus) in LSA framework. This matrix tells how many times a word occurs in a certain document. In other words, this matrix ignores the word order in the sentence. We propose several word co-occurrence matrices that keep the word order. By applying LSA to these matrices, words in the vocabulary are projected to a continues vector space according to their position in the sentences. To support these matrices, we define a context dependent class (CDC) based ngram. Unlike traditional class-based n-gram LM, CDC LM distinguishes classes according to their context in the sentences. Experiments on Wall Street Journal (WSJ) corpus show that the word co-occurrence matrix works 3.62%-12.72% better in terms of perplexity than word-document matrix. Furthermore, the CDC improves the performance and achieves better perplexity than the traditional class-based n-gram LM based on LSA. When the model is linearly interpolated with the word-based 3-gram, it gives improvements about 2.01% for 3-gram model and 9.47% for 4-gram model on relative perplexity against a standard word-based 3-gram LM. During the past few years, researchers have tried to incorporate long-range dependencies into statistical word-based n-gram LMs. One of these long-range dependencies is topic. Unlike words, topic is unobservable. Thus, it is required to find the meanings behind the words to get into the topic. As the second part of this research, we proposed a new approach for a topic-dependent LM called topic dependent class (TDC) based n-gram, where the topic is decided in

Read the paper · More papers on PaperTik