Robust speech recognition in car environments
Makoto Shozakai, Satoshi Nakamura, Kiyohiro Shikano · 2002
A user-friendly speech interface in a car cabin is highly needed for safety reasons. This paper describes a robust speech recognition method that can cope with additive noise and multiplicative distortions. A known additive noise, a source signal of which is available, might be canceled by NLMS-VAD (normalized least mean squares with frame-wise voice activity detection). On the other hand, an unknown additive noise, a source signal of which is not available, is suppressed with CSS (continuous spectral subtraction). Furthermore, various multiplicative distortions are simultaneously compensated with E-CMN (exact cepstrum mean normalization) which is speaker dependent/environment-dependent CMN for speech/non-speech. Evaluation results of the proposed method for car cabin environments are finally described.