Hybrid coding of speech at low bit-rate
E. Shlomot, A. Gersho · 1998
Low bit-rate coding of speech signals plays an important role in digital speech communication systems, enabling speech transmission over power-limited or band-limited wireline and wireless networks, and the recording of speech on space-limited storage media. harmonic coding of speech uses the harmonic spectral structure of voiced speech and the noise characteristics of unvoiced speech for modeling and coding. In this work we propose a new model for high-quality low bit-rate coding of speech, which combines a harmonic model for voiced speech, a noise model for stationary unvoiced speech, and a waveform model for transition speech, and we call it “hybrid coding.” Several design problems and implementation issues are addressed for hybrid modeling and coding of speech. A crucial issue for the hybrid model is the operation of a robust speech classifier to distinguish between the different speech classes. For the coding of voiced speech, the pitch and harmonic bandwidth (voicing) parameters must be reliably estimated. A unique problem in hybrid coding is to find effective waveform. modeling and coding methods for transition speech. We investigated and solved these problems, and designed a hybrid coder based on our solutions. Our speech classifier was implemented by a neural network, which provided a robust speech classification for clean speech and for speech in the presence of background noise. Harmonic matching parameters were used for speech classification, pitch detection, and harmonic bandwidth estimation. Transition speech was modeled and coded by a multi-pulse waveform matching approach which, to some extent, was also suitable for the modeling of both voiced and unvoiced speech, and provided model “overlap” for coding robustness. Phase synchronization was used for signal alignment when switching between the harmonic model of voiced speech and the waveform model of transition speech. Harmonic spectral quantization requires a combined scheme for dimension conversion and vector quantization. We demonstrated in this work that several known schemes for dimension conversion can be unified by the non-square linear transform approach. We showed that linear transforms are particularly suitable for weighted vector quantization of the harmonic spectral samples of the residual signal, and we addressed the problem of low-complexity perceptually-weighted linear dimension conversion and vector quantization. We applied our solutions to the hybrid coding approach and implemented a 4 kbps hybrid speech coder. Subjective quality tests demonstrated that our 4 kbps hybrid coder is capable of delivering a high-quality synthesized speech, for clean speech and for speech in the presence of stationary room noise. (Abstract shortened by UMI.)