Continuous speech recognition for the TIMIT database using neural networks
F. Fallside, Helmut Lucke, T.P. Marsland, Peter O’Shea, Mallory Owen, Richard W. Prager, A.J. Robinson, N.H. Russell · International Conference on Acoustics, Speech, and Signal Processing · 2002
Four types of neural networks which have previously been established for speech recognition and tested on a small, seven-speaker, 100-sentence database are applied to the TIMIT database. The networks are a recurrent network phoneme recognizer, a modified Kanerva model morph recognizer, a compositional representation phoneme-to-word recognizer, and a modified Kanerva model morph-to-word recognizer. The major result is for the recurrent net, giving a phoneme recognition accuracy of 57% from the si and sx sentences. The Kanerva morph recognizer achieves 66.2% accuracy for a small subset of the sa and sx sentences. The results for the word recognizers are incomplete.>