On the Use of Forward Temporal Masking and Cumulative Distribution Mapping for Noisy Speech Recognition

Eric H. C. Choi, Julien Epps · 2005

Robustness in the presence of various types and levels of environmental noise remains an important issue for automatic speech recognition (ASR) systems. This paper describes a new noise-robust ASR front-end that employs a functional model of forward temporal masking combined with cumulative distribution mapping based on MFCC's with c0. Recognition experiments on the Aurora II connected digits database reveal that the proposed front-end achieves an average digit recognition accuracy of 83.24% for a model set trained from clean data and 90.32% for a model set trained from data with multiple noise conditions. Compared with the ETSI standard Mel-cepstral front-end, the proposed front-end obtains a relative error reduction of around 57% for the clean model set and 21% for the multi-condition model set.

Read the paper · More papers on PaperTik