Phoneme recognition using a time-sliced recurrent recognizer
Ingrid Kirschning, Hideto Tomabechi · 2002
This paper presents a new method for phoneme recognition using neural networks, the time-sliced recurrent recognizer (TSRR). In this method we employ Elman's recurrent network with error-backpropagation, adding an extra group of units that are trained to give a specific representation of each phoneme while it is recognizing it. The purpose of this architecture is to obtain an immediate hypothesis of the speech input without having to pre-label each phoneme or separate them before the input. The input signal is divided into time-slices which are recognized in a linear sequential fashion. The generated hypothesis is shown in the extra group of units at the same moment the time-slices are passed through the network and being recognized as a certain phoneme. Thus the TSRR is capable of recognizing the phonemes in real-time without discriminatory learning.>