A comparative study of signal representations and classification techniques for speech recognition
Hong C. Leung, Benjamin Chigier, James Glass · IEEE International Conference on Acoustics Speech and Signal Processing · 1993
The authors investigate the interactions of two important sets of techniques in speech recognition: signal representation and classification. In addition, in order to quantify the effect of the telephone network, experiments are performed on both wideband and telephone-quality speech. The spectral and cepstral signal processing techniques studied fall into a few major categories based on Fourier analyses, linear prediction, and auditory processing. The classification techniques examined are Gaussian, mixture Gaussians, and the multilayer perceptron (MLP). Results indicate that the MLP consistently produces lower error rates than the other two classifiers. When averaged across all three classifiers, the Bark auditory spectral coefficients (BASC) produce the lowest phonetic classification error rates. When evaluated in a stochastic segment framework using the MLP, BASC also produces the lowest word error rate.>