Document-level optimization in speech recognition

Rie Nakazato, Kugatsu Sadamitsu, Mikio Yamamoto · The Journal of the Acoustical Society of America · 2006

Most ASR systems optimize scores at the sentence level because language models are designed to assign a meaningful probability to a sentence, but not to a document. However, in the last half-decade, NIPS people have specifically examined aspect-based global language models and developed generative text models such as the latent Dirichlet allocation (LDA), which can assign a meaningful probability to an entire document. We investigated a system to recognize all sentences in a read document, maximizing a global score including the document probability, in addition to assigning acoustic and local language probabilities. Document level optimization using global scores engenders a combinatorial problem in its implementation. Therefore, we used the rescoring framework with N-best recognition results for each sentence in a read document and the hill-climbing method to search for approximately optimal recognition results. In this study, we compare mixture of unigrams (MU), Dirichlet mixtures (DM), and LDAs as generative text models in experiments using read document speech data in Japanese Newspaper Article Sentences (JNAS) and academic presentation speech data in The Corpus of Spontaneous Japanese (CSJ).

Read the paper · More papers on PaperTik