Waveform quantization of speech using Gaussian mixture models

Johan Samuelsson · 2004

Waveform quantization of speech using Gaussian mixture models (GMM) is proposed. GMM are trained directly on the speech waveform, and high dimensional vector quantizers (VQ) that efficiently exploit the redundancy are constructed based on the GMM parameters. Two types of GMM are studied. The complexity of the scheme is independent of the rate, and the rate can be changed without retraining the VQ. A shape-gain structure improves performance and robustness. Pre- and post-processing using spectral amplitude warping further improves perceptual quality. A 32-dimensional VQ operating at 2 bits/sample reproduces speech sampled at 8 kHz with a PESQ score of 4.2.

Read the paper · More papers on PaperTik