Recurrent neural network speech predictor based on dynamical systems approach

Ekrem Varoğlu, Kadri Hacıoğlu · IEE Proceedings - Vision Image and Signal Processing · 2000

A nonlinear predictive model of speech, based on the method of time delay reconstruction, is presented and approximated using a fully connected recurrent neural network (RNN) followed by a linear combiner. This novel combination of the well established approaches for speech analysis and synthesis is compared with traditional techniques within a unified framework to illustrate the advantages of using an RNN. Extensive simulations are carried out to justify the expectations. Specifically, the network's robustness to the selection of reconstruction parameters, the embedding time delay and dimension, is intuitively discussed and experimentally verified. In all cases, the proposed network was found to be a good solution for both prediction and synthesis.

Read the paper · More papers on PaperTik