Performance Evaluation of Front-end Processing for Speech Recognition Systems
Octavian Cheng, Waleed Habib Abdulla, Zoran A. Salcic · 2005
Conventional speech feature extraction front-end algorithms suffer severe performance degradation in noisy environment, especially when there is a noise level mismatch between the training and testing environments. It is necessary to search for new feature extraction algorithms, which perform better than the convention methods in adverse conditions. In our literature survey, two recently developed algorithms are found to have better performances than the conventional algorithms. They are Gammatone Cepstral Coefficients (GTCC) and Zero-Crossings with Peak Amplitude (ZCPA). Two sets of experiments are conducted to test their performances against the conventional methods, which include Linear Prediction Cepstral Coefficients (LPCC), Perceptual Linear Prediction Coefficients (PLP) and Mel Frequency Cepstral Coefficients (MFCC). The first set of experiments is a pilot study, which involves recognising speaker-dependent isolated numeric digits. The second set of experiments is a more formal study, which includes HMM-based speaker-independent continuous speech recognition experiments using an accredited speech database called TIMIT. In