A speech recognition system consisting of auditory feature extracting cells and velocity-controlled delay-lines. II. Recognition module

K. Yamauchi, Masakatsu Fukuda, Kunihiko Fukushima · 2005

A neural network model for speech recognition is proposed, based on neurophysiological findings of the auditory system. The first stage of the system is a feature extracting module. The extracted features are sent to the next stage, the recognition module. The recognition module consists of three blocks of multilayered networks: the upper, middle, and lower blocks. Each block is a neocognitron-like network whose first layer consists of velocity-controlled delay-lines. The propagation velocities of the delay-lines of the upper and lower blocks are faster and slower, respectively, than those of the middle one. The propagation velocities of these three delay-lines are variable but the ratio of the velocities is always kept constant. This system controls the velocity of all of the delay-lines in such a way that the durations of features on the delay-line of the middle network are the same as the duration of features of a training pattern. This is accomplished by comparing the outputs of the upper and lower blocks. In the simulation, the system was trained using some Japanese words.

Read the paper · More papers on PaperTik