Speech analysis and recognition using interval statistics generated from a composite auditory model

Hamid Sheikhzadeh, Li Deng · IEEE Transactions on Speech and Audio Processing · 1998

A modeling approach to auditory speech analysis and recognition is proposed and evaluated, where a composite auditory model is used to generate parallel sets of auditory-nerve instantaneous firing rates (IFRs) along the spatial dimension, followed by a processing stage that constructs from the IFRs the interval statistics in a form called the interpeak interval histogram (IPIH). A speech preprocessor is designed that performs transformation on the auditory IPIHs and interfaces the IPIH-based auditory representation with a hidden Markov model-based (HMM-based) speech recognizer. The results demonstrate that the new preprocessor consistently outperforms the conventional mel frequency cepstral coefficient-based (MFCC-based) preprocessor for the signal-to-noise ratio (SNR) level up to at least 16 dB.

Read the paper · More papers on PaperTik