A comparative study of five spectral representations for speaker-independent phonetic recognition

J. Creekmore, Mark Fanty, Ronald A. Cole · 2002

The authors describe a comparative study of five spectral representations for speaker-independent phonetic recognition using the TIMIT database. A feedforward network was trained to classify 20-ms frames of speech as one of 39 phonetic classes derived from the TIMIT database. The five representations investigated include the discrete Fourier transform, three representations based on conventional linear predictive coding (LPC), and the cepstral coefficients derived from perceptual linear predictive (PLP) analysis. The PLP cepstral coefficients outperformed the other representations on the task of assigning the correct phonetic label to individual time frames. It is shown that phonetic context can be exploited by providing spectral information before and after the frame to be classified. The effect of the training set size and distribution is also examined.>

Read the paper · More papers on PaperTik