Discrete wavelet transform techniques in speech processing

Johnson Ihyeh Agbinya · 2002

The trend towards real-time, low-bit-rate speech coders dictates current research efforts in speech compression. A method being evaluated uses wavelets for speech analysis and synthesis. Distinguishing between voiced and unvoiced speech, determining pitch, and methods for choosing optimum wavelets for speech compression are discussed. We observe that wavelets concentrate speech energy into bands which differentiate between voiced or unvoiced speech. Optimum wavelets are selected based on energy conservation properties in the approximation part of the wavelet coefficients. It is shown that the Battle-Lemarie wavelet concentrates more than 97.5% of the signal energy into the approximation part of the coefficients followed closely by the Daubechies D20, D12, D10 or D8 wavelets. The Haar wavelets are the worst. Listening tests show that the Daubechies 10 preserves perceptual information better than other Daubechies wavelets and, indeed, a host of other orthogonal wavelets. Pitch periods and evolution can be identified from contour plots of coefficients obtained at several scales.

Read the paper · More papers on PaperTik