Noise robust speech recognition based on spatial subtraction array
Yasutaka Ohashi, Tsuyoki Nishikawa, Hiroshi Saruwatari, Akinobu Lee, Kiyohiro Shikano · 2005
Summary form only given. We propose a spatial subtraction array (SSA) and known noise superimposition to achieve a robust hands-free speech recognition under noisy environments. In the proposed SSA, noise reduction is achieved by subtracting the estimated noise power spectrum from the target speech power spectrum to be enhanced in the mel-scale filter bank domain. This offers a realization of error-robust spatial spectral subtraction with few computational complexities. In addition, we introduce known noise superimposition techniques in the mel-scale filter bank domain, and utilize the matched acoustic model for the known noise. This can compensate the acoustic model mismatch and mask the residual noise component in SSA. The experimental results obtained under a real environment reveal that word accuracy of the proposed method is greater than that of the conventional method even when the target user moves between -10 and +10 degrees around the microphone array.