Single Channel Noise Reduction for Hands Free Operation in Automotive Environments

S. Schmitt, Malte Sandrock, Jochen Cronemeyer · Journal of the Audio Engineering Society · 2002

The objective of this work is to establish a single channel noise reduction algorithm for speech enhancement integrated in DSP systems. The main emphasis is on spectral weighting. The chosen algorithm is based on a Minimum Mean Square Error Log Spectral Amplitude approach. One of the crucial tasks for good results, i.e. natural and intelligible speech in combination with well attenuated noise and low spectral distortion, is a balanced estimation and weighting of the noise magnitude spectrum. INTRODUCTION In modern hands free speech communication environments often occurs the situation that the speech signal is superposed by background noise (see Figure 1). This is particular the case if the speaker is not located as close as possible to the microphone. The speech signal intensity decreases with growing distance to the microphone. It is even possible that background noise sources are captured at a higher level than the speech signal. The noise distorts the speech and words are hardly intelligible. In order to improve the intelligibility and reduce the listeners (FES) stress by increasing the signal to noise ratio a noise reduction or in a wider sense – an also called speech enhancement algorithm is applied. The objective of this work is to establish a model of a single channel speech enhancement algorithm using MATLAB as a base for a DSP software implementation. The main emphasis is on spectral weighting. The chosen algorithm with the best results is based on a Minimum Mean-Square Error Log-Spectral Amplitude (MMSE-LSA) approach. NOISE REDUCTION PRINCIPLES The requirements of a noise reduction system for speech enhancement are: • Intelligibility and naturalness of the enhanced signal • Improvement of signal-to-noise ratio • Short signal delay • Computational simplicity The quality of the enhanced signal is a diverse issue, it may be characterised by the terms intelligibility and naturalness. There are several methods for performing noise reduction, but all can be regarded as a kind of filtering. In our application speech and noise are mixed to one signal channel. They reside in the same frequency band and may have similar correlation properties. Consequently the filtering will inevitably have an effect on both the speech and the noise. Therefore it is a very challenging task to distinguish between them. I.e. speech components can be detected as noise and thus will be suppressed as well. Especially fricatives and plusives are attenuated due to their noise-like properties. Furthermore the SCHMITT ET AL. SINGLE CHANNEL NOISE REDUCTION AES 112 CONVENTION, MUNICH, GERMANY, 2002 MAY 11–14 2 residual noise characteristics should preserve the characteristics of the background noise in the recording environment. Typical single channel noise reduction algorithms add a synthetic noise, also called ‘Musical Noise’, which sounds artificial and has a disturbing effect on the listener. Single channel noise reduction algorithms are based on the fact that the statistical properties of speech are only stationary over short periods of time whereas the noise often can be assumed to be stationary over much longer periods. Another aim for the algorithm design is the limitation of the signal delay because of its annoying effect in dialog situations. The noise reduction algorithms can be split into two groups: time domain algorithms and those utilising some kind of transform, e.g. Fourier Transform. Whereas the filter calculation for time domain solutions generally relies on the usage of correlation estimates, there is a large variety of algorithms operating in the frequency domain. Noise reduction in frequency domain The fundamental concept of a frequency domain solution is spectral weighting and block processing. The architecture of such a system is presented in Figure 2. It consists of three major components: • the analysis/synthesis framework for time domain / frequency domain transformation • the noise estimation • the weighting function. In a typical hands free situation (Fig. 1) the recorded time domain signal x(k) is composed of the superposition of speech s(k) and noise n(k): ) ( ) ( ) ( k n k s k x + = (equ. 1) The basic idea of spectral subtraction is to estimate the noise spectrum Nest(n,Ωi) and to subtract it from the observed signal spectrum X(n,Ωi): Y(n,Ωi) = Sest(n,Ωi) = X(n,Ωi) – Nest(n,Ωi) (equ. 2) where n designates the current frame and Ωi the frequency bin. If the noise estimation equals the disturbing noise spectrum the output signal spectrum Y(n,Ωi), also designated as speech estimation Sest(n,Ωi), will be very similar to the noiseless speech spectrum S(n,Ωi). Because simple spectral subtraction shows limited performance in manner of speech quality, equ. 2 is converted to ) , ( ) , ( ) , ( ) , ( 1 ) , ( ) , (

Read the paper · More papers on PaperTik