On the training strategies of neural networks for speech recognition
Fikret Sadik Gürgen, K. Aikawa, Kiyohiro Shikano · 2003
The authors investigate how to introduce invariant features to speech recognition neural networks using conventional back propagation (BP), K-neighbor interpolation training (KNIT) with a number of time-shifted examples (TSEs) of the same training sample. The TSEs are employed for training of a multilayer perceptron (MLP) and a time-delay neural network (TDNN) structure to enrich the training sample set covering a larger area of phoneme sample space. Speaker-dependent phoneme recognition experiments were performed. The advantages and disadvantages of using time-shifted examples of a training sample for a MLP and a TDNN structure and a BP and a KNIT algorithm are discussed.>