Source and channel coding for speech transmission and remote speech recognition
Alexis Bernard, Abeer A. Alwan · 2002
This dissertation addresses the issue of designing source and channel coding techniques for two types of speech processing applications: speech transmission and remote speech recognition. In the first part, adaptive multi-rate (AMR) speech transmission systems that switch between operating modes depending on channel conditions are presented. We address the design of such an adaptive scheme using variable bit rate embedded source encoders and rate-compatible channel coders providing unequal error protection. A novel technique, the rate-compatible punctured trellis code (RCPT) for obtaining unequal error protection via progressive puncturing of symbols in a trellis, is presented and compared with the rate-compatible punctured convolutional code with and without bit-interleaved coded modulation. The perceptually-based speech coder proposed displays a wide range of bit error sensitivities, and is used in combination with rate-compatible punctured channel codes providing adequate levels of protection. The resulting system operates over a wide range of channel conditions with graceful performance degradation as the channel signal-to-noise ratio decreases. In the second part, we present a framework for developing source coding, channel coding, channel decoding, and frame erasure concealment techniques adapted for remote speech recognition applications. It is shown that speech recognition, as opposed to speech coding, is more sensitive to channel errors than channel erasures. Appropriate channel coding design criteria are determined. For channel decoding, we introduce a novel technique for combining soft decision decoding with error detection. The technique outperforms the often used hard decision strategy. In addition, frame erasure concealment techniques are used at the decoder to deal with unreliable frames. At the recognition stage, we present a technique to modify the recognition engine to take into account the time-varying reliability of the decoded feature after channel transmission. The resulting engine, referred to as weighted Viterbi recognition (WVR), further improves recognition accuracy. Together, source coding, channel coding and the modified recognition engine are shown to provide good recognition accuracy over a wide range of communication channels at very low bit rates.