A recurrent time-delay neural network for improved phoneme recognition
F. Greco, Andrea Paoloni, Giacomo Ravaioli · 1991
The authors propose a modification to the structure of the time-delay neural network (TDNN), obtained through feedback at the first-hidden layer level. The experiment carried out with the new model, called RTDNN (recurrent TDNN), consists of the classification of the unvoiced plosive phonemes. These were extracted from an initial and intermediate position in a list of the most common Italian words, uttered by a male speaker, thus obtaining 250 tokens per phoneme. The training was carried out through a modified variant of back propagation, known as BPS (back propagation for sequences), using half of the tokens for learning and the remaining for the test. The error rate trend thus obtained shows a 27% decrease in a particular range of the magnitude of feedback, with values ranging from 5% for the original TDNN model with no feedback to 3.6% for the proposed RTDNN model.>