Frequency-time-shift-invariant time-delay neural networks for robust continuous speech recognition

H. Sawai · 1991

The authors propose neural network (NN) architectures for robust speaker-independent, continuous speech recognition. One architecture is the frequency-time-shift-invariant time-delay neural network (FTDNN). Another architecture is based on windowing each layer of the NN with local time-frequency windows. This architecture makes it possible for the NN to capture global features from the upper layers as well as precise local features from the lower layers. Recognition experiments on easily confused phonemes were performed using /b/, /d/, /g/, /m/, /n/, and /N/ (syllabic nasal) phoneme tokens to verify robustness to variations of speech. Performance results for the different architectures are presented.>

Read the paper · More papers on PaperTik