An acoustic front-end using warped frequency and temporal resolutions
B.T. Lilly, Kuldip K. Paliwal · 2002
Typically, the power spectrum of a speech frame used in speech recognition is estimated for a fixed length window using the fast Fourier transform. Each frequency component represented in this power spectrum is an estimate over that speech frame. The power spectrum calculated in this way has a constant time and frequency resolution. An example of this type of front-end is the LPC-derived cepstral front-end commonly used is recognition systems today. The acoustic front-end presented in this paper employs both a warped frequency and temporal resolutions. We show that a front-end that utilises both warping functions, outperforms a front-end that employs only a warped frequency scale. We also show that this new front-end is unsuitable for noisy conditions.