Connectionist techniques for speaker-independent recognition of isolated utterances

Michael A. Franzini · The Journal of the Acoustical Society of America · 1988

The performance of previous speech-recognition systems has been limited to a large extent by the incomplete structural theories of speech upon which the systems were based. The work reported here suggests that connectionist learning procedures provide an effective method for generating speaker-independent recognition systems without the need for any a priori model of human speech production. The backpropagation learning algorithm is applied to the problem of isolated-utterance speech recognition, and the resulting recognition rates approach those of the best conventional systems. Several studies were peformed using computer simulations of various networks trained by backpropagation. With 1 s of digitized speech as input, the networks had to generate as output the appropriate labels, which in these studies were letters of the alphabet. The accuracy for speaker-independent recognition of the whole alphabet reached 89%, which is higher than the accuracies achieved by several more conventional recognizers using the same database [K. F. Lee, Carnegie Mellon Univ. Tech Rep. 85–181 (1985)]. Recognition rates for speaker-dependent recognition of the whole alphabet reached 99% and, for speaker-independent recognition of confusable letter sets such as b-p-e-v-d, the networks achieved 94%. These studies demonstrate that networks with simple task-independent learning procedures can perform as well as systems that explicitly implement procedures such as dynamic time warping and vector quantization of inputs.

Read the paper · More papers on PaperTik