A Novel Acoustic Feature Extraction Algorithm Based on Root Cepstrum Coefficients and CCBC for Robust Speech Recognition

Xu Wang, Zhiyan Han · 2008

Studies have shown that depending on speaker task and environmental conditions, recognizers are sensitive to noisy stressful environments. The focus of this study is to achieve robust recognition in diverse environmental conditions through extracting robust features. Central to the technique is Root Cepstrum Coefficients (RCC) method, instead of logarithm amplitude spectrum and discrete cosine transform of the conventional Mel Frequency Cepstral Coefficients (MFCC), but also using Two-dimensional Root Cepstrum Coefficients (TDRCC). This feature is called TDRCC-MFCC. And then, we consider incorporating Canonical Correlation Based Compensation (CCBC) to cope with the mismatch between training and test set. The mismatch between training and test conditions can be simply clustered into three classes: differences of speakers, changes of recording channel and effects of noisy environment. We evaluate the technique using Back-Propagation Neural Networks (BPNN) on two different tasks: one is in-car speech recognition task, another is different SNR speech recognition. The experimental results show that the novel feature has very good robustness and effectiveness relative to MFCC feature and the CCBC algorithm can make speech recognition system greatly robust to all three kinds of mismatch between training set and test set.

Read the paper · More papers on PaperTik