Lombard effect compensation and noise suppression for noisy Lombard speech recognition

Sang-Mun Chi, Yung‐Hwan Oh · 2002

The performance of a speech recognition system degrades rapidly in the presence of ambient noise. To reduce the degradation, a degradation model is proposed which represents the spectral changes in a speech signal uttered in a noisy environment. The model uses frequency warping and amplitude scaling of each frequency band to simulate the variations of formant location, formant bandwidth, pitch, spectral tilt and energy in each frequency band by the Lombard effect. Another Lombard effect-the variation of overall vocal intensity-is represented by a multiplicative constant term depending on the spectral magnitude of the input speech. The noise contamination is represented by an additive term in the frequency domain. According to this degradation model, the cepstral vector of clean speech is estimated from that of noisy-Lombard speech using spectral subtraction, spectral magnitude normalization, band-pass filtering in the Lin-Log spectral domain, and multiple linear transformations. Noisy Lombard speech data is collected by simulating noisy environments using noises from automobiles, an exhibition hall, telephone booths in downtown crowded streets, and computer rooms. The proposed method significantly reduces error rates in the recognition of 50 Korean words. For example, the recognition rate is 95.91% with this method and 79.68% without this method at an SNR (signal-to-noise ratio) 10 dB.

Read the paper · More papers on PaperTik