Speaker-independent word recognition using a neural prediction model

K.-I. Iso, Takao Watanabe · International Conference on Acoustics, Speech, and Signal Processing · 2002

A speech recognition model called the neural prediction model (NPM) is proposed. The model uses a sequence of multilayer perceptrons (MLPs) as a separate nonlinear predictor for each class. It is designed to represent temporal structures of speech patterns as recognition cues. In particular, temporal correlation in successive feature vectors of a speech pattern is represented in the mappings formed as MLP input-output relations. Temporal distortion of speech is efficiently normalized by a dynamic-programming technique. Recognition and training algorithms are presented based on the combination of dynamic-programming and back-propagation techniques. Evaluation experiments were conducted using ten-digit vocabulary samples uttered by 107 speakers. A 99.8% recognition accuracy was obtained. This suggests that the model is effective for speaker-independent speech recognition.>

Read the paper · More papers on PaperTik