A Neural Network for Classification of Spoken Digits
Thomas M. English, Lois Boggess · 1990
In most applications of neural networks to speech recognition, the task of the network is to emit exactly one classification in response to a sequence of one or more inputs. This paper describes a neural net that emits sequences of classifications in response to sequences of inputs. At fixed intervals, a representation of the speech spectrum is input to the network. When the network detects the completion of an utterance of a digit, it identifies the digit. At all other times, the network emits a null classification. This behavior is achieved through three stages of nonlinear transduction. Speech spectra are input to a Kohonen map, the outputs of which are input to a recurrent subnetwork (i.e., one with feedback connections). The outputs of the recurrent subnet are classified by a feed-forward subnet. The recurrent subnet, which implements the requisite memory of previous inputs, has no trainable parameters. Supervised back-propagation training of the feed-forward subnet follows unsupervised training of the Kohonen map. In an evaluation of eight variants of the network, random strings of digits were read by a single talker. Each network was trained with 385 strings (1540 words), and was tested with 140 strings (560 words). The best variant classified the training and test utterances with word-error rates of 0% and 3.8%, respectively. Four other variants achieved word-error rates not exceeding 3.9% in testing.