Predictive, hierarchical, and transform vector quantization for speech coding
Pao‐Chi Chang · 1986
Vector Quantization has proved useful for speech and image coding applications. Shannon theory states that vector quantization can achieve nearly optimal performance when the vector dimension is sufficiently large. Unfortunately, the codebook size grows exponentially with vector dimension and hence the search complexity also grows exponentially. For this reason recent research has focused on various structures which either yield better performance or need less complexity than the original vector quantizer for a given rate. Three structures and design algorithms are presented in this thesis: (1) Predictive vector quantizers which are vector extensions of a predictive quantizer or a differential pulse code modulation (DPCM) system. Two gradient algorithms modified from adaptive filtering techniques for designing such a system are developed: the steepest descent algorithm and the stochastic gradient algorithm. (2) Hierarchical vector quantizers which implement vector quantizer encoders by table lookups rather than by minimum distortion searches so as to eliminate arithmetic computations in on-line coding. Hierarchical structures are used to preserve manageable table sizes for large dimension vector quantizers. (3) Fourier transform vector quantizers which vector quantize Fourier transformed data instead of time domain waveforms. A product code structure and a properly chosen bit allocation among codebooks are used to quantize the transformed coefficients efficiently. Both real-imaginary and magnitude-phase systems with corresponding distortion measures are considered. Both systems yield good performance for a given complexity in comparison with waveform vector quantizers in speech coding. Code designs and tests are simulated for the above systems with sampled speech waveforms.