Wavelet-based perceptual speech enhancement using adaptive threshold estimation
Essa Jafer, Abdulhussain E. Mahdi · 2003
A new speech enhancement system, which is based on a timefrequency adaptive wavelet soft thresholding, is presented in this paper. The system utilises a Bark-scaled wavelet packet decomposition integrated into a modified Weiner filtering technique using a novel threshold estimation method based on a magnitude decision-directed approach. First, a Bark-Scaled wavelet packet transform is used to decompose the speech signal into critical bands. Threshold estimation is then performed for each wavelet band according to an adaptive noise level-tracking algorithm. Finally, the speech is estimated by incorporating the computed threshold into a Wiener filtering process, using the magnitude decision-directed approach. The proposed speech enhancement technique has been tested with various stationary and non-stationary noise cases. Reported results show that the system is capable of a high-level of noise suppression while preserving the intelligibility and naturalness of the speech. With the rapid developments in voice communication systems, speech enhancement using a single channel has become an active and important research area. During the last few decades, various approaches to enhance the speech quality by reducing the noise have been proposed. The most widely used methods are those based on spectral subtraction and Wiener filtering and their variants. Although most of these methods have been shown to provide good speech quality particularly in terms of improved signal-to-noise ratio (SNR), they often suffer from an annoying signal distortion caused by a residual effect known as musical noise. In an attempt to reduce this drawback, the use of a human auditory model has recently been proposed in subtractive-type enhancement techniques