Lattice Parsing for Speech Recognition
Chappelier, Jean-Cédric, Martin Rajman, Aragües, Ramon, Antoine Rozenknop · Infoscience (Ecole Polytechnique Fédérale de Lausanne) · 1999
A lot of work remains to be done in the domain of a better integration of speech recognition and language processing systems. This paper gives an overview of several strategies for integrating linguistic models into speech understanding systems and investigates several ways of producing sets of hypotheses that include more “semantic ” variability than usual language models. The main goal is to present and demonstrate by actual experiments that sequential coupling may be efficiently achieved by word-lattice syntactic analyzers, efficiently parsing the huge number of hypothesis (i.e. possible sentences) contained in the lattice produced by the speech recognizer. 1. Motivations The past decade has seen significant progress in speech recognition technology: word (recognition) error rates continue to drop by a factor of 2 every two years (Rabiner et al., 1996) and high performance systems are now becoming available. Several factors have contributed to this rapid progress: Generalisation and continuous improvements of the powerful Hidden Markov Model (HMM); Better language models and powerful algorithms allowing their integration in speech recognition systems; Production of large speech corpora allowing researchers to optimize parameters of the recognizers in a statistically meaningful way; Establishment of standards for performance evaluation and advances in computer technology. However, speech recognition remains a difficult problem, due to the large variability associated with the input signal it considers. In particular a lot of work is to be done towards a better integration of linguistic models into continuous speech recognition systems, although several ways have already been studied: statistical models (bigrams and trigrams), stochastic or deterministic finite state automata (FSA), and context-free grammars (CFGs). These approaches are indeed mainly designed to reduce the overall word-error-rate and are not necessarily appropriate for higher level knowledge post-processing. The current paper investigates several ways of producing sets of hypotheses 1 that include more “semantic ” variability, becoming therefore more appropriate for higher-level linguistic post-processing.