Speech spectrum quantization using Gaussian mixture models and multi-dimensional companding
A.D. Subramaniam, William R. Gardner, Bhaskar D. Rao · 2003
A low complexity, high quality, quantization scheme using Gaussian mixture models and multi-dimensional companding is presented for speech spectrum quantization. As in past work (Subramaniam and Rao 2000), the probability density function (pdf) of the source is modeled as a Gaussian mixture model whose parameters are estimated using the EM algorithm. The novelty of the approach lies in the development of a more efficient mapping from the density to the codebook vectors. The individual Gaussian component densities are quantized using an optimal multidimensional compressor function followed by a lattice vector quantizer (LVQ). For the purpose of speech spectrum quantization, the E/sub 8/ lattice is used for lattice quantization. Low complexity encoding and index assignment algorithms are presented and the performance is compared and shown to be better than that of PDF optimized parametric VQ.