Exploiting high-level knowledge resources for speech recognition with applications to interactive voice response systems

DAN I. MOLDOVAN, Mithun Balakrishna · 2007

This dissertation proposes a novel methodology to improve the performance of a Large Vocabulary Continuous Speech Recognizer (LVCSR) by modeling several high-level knowledge resources into an n-best list re-ranking mechanism, and its use in a novel Interactive Voice Response (IVR) grammar creation/tuning application. First, this dissertation focuses on the identification and formulation of several novel, additional, domain-independent knowledge resources into a re-ranking mechanism. We illustrate the extent of improvements obtainable by efficiently exploiting: phonetic knowledge presented by an additional classifier, lexical knowledge presented by a dominant word-based lexical scoring mechanism, syntactic knowledge presented by a statistical, immediate-head-based syntactic parser and semantic knowledge presented by a predicate argument coverage algorithm using the basic semantic propositions of a language. We achieve 2.9%, 6.6%, and 1.1% absolute Word Error Rate (WER) improvements for spontaneous, directed-dialog and spoken-questions evaluation sets respectively. Secondly, we improve WER for specific domains by combining domain-independent knowledge with automatically extractable domain-dependent resources. To model domain-dependent knowledge, we propose a methodology to automatically generate Statistical Language Models (SLMs) for specific dialog states, and a semantic category classification strength extraction mechanism. On a directed-dialog evaluation set, we obtain 4.4% absolute WER reduction (using automatically generated SLMs only) over the best manual grammar based system while the Semantic Error Rate (SemER) is higher by 1%. Combining the domain-independent and dependent knowledge resources, we achieve 10.1% absolute WER improvement for a directed-dialog test set, over the baseline LVCSR. The heart of this dissertation not only lies in the task of selecting and modeling key information resources but also on combining them efficiently. Hence, we explore using minimum error rate training to optimally assign knowledge resource weights by directly minimizing the WER on a development set. Finally, we present a novel IVR grammar creation/tuning application and illustrate the importance of the re-ranking mechanism in this framework. The automatic and efficient grammar creation/tuning application architecture contains a semantic categorizer, a statistical filtering mechanism, and a mechanism to raise an alarm condition in case of large inflow of errors. Results show improved IVR performance that closely matches (and sometimes beats) the manual tuning performance.

Read the paper · More papers on PaperTik