Feature extraction from a cochlear model representation: A novel supervised/unsupervised neural network hybrid for speech recognition.

Gary N. Tajchman, Nathan Intrator · The Journal of the Acoustical Society of America · 1992

A novel supervised/unsupervised hybrid neural network training algorithm [Intrator, ‘‘Combining exploratory projection pursuit with projection pursuit regression,’’ preprint (1992)] is applied to a speech recognition task. The input representation for the neural network is produced by Richard Lyon’s cochlear model as implemented by Slaney [M. Slaney, ‘‘Lyon’s Cochlear Model,’’ Apple Tech. Report ♯13, Apple Comput., Inc., Cupertino, CA 95014]. This detailed, high-dimensional representation is projected through the units of the hidden layer of a back-propagation-like architecture into a much lower dimensional space (spanned by the number of hidden units). This space is constructed by optimizing a combination of the familiar MSE minimization and an unsupervised measure of the goodness of the projections. The goodness measure is based on the multimodality of the projected distributions. Both constraints are powerful and useful when used alone. However, when they are combined, the resulting network has the potential to learn and generalize much more robustly than either alone. This technique was applied to the classification of 16 classes of stressed vowels extracted from the TIMIT Acoustic-Phonetic Continuous Speech Corpus [NIST Speech Disc 1-1.1 (October 1990)]. Comparisons between the hybrid method and plain back propagation are discussed.

Read the paper · More papers on PaperTik