Visual representation of the speech trace using neural networks
P. Gomez, V. Rodellar, Agustín Álvarez-Marquina, N. Mayo, F. Rubio, Victor M. Garca Nieto, M.M. Perez · 2002
Through the present paper, a methodology to create Visual Representations of Speech for Speech Perception Enhancement Applications, is presented, based on the use of Time-Delay Neural Networks. The advantages of using Neural Networks for such purposes, come from a lower computational cost, and from an easier DSP or VLSI implementation. On the other hand, the main inconvenient found in using this technique, is the need for training to each specific Speaker. This requirement may be relaxed if proper normalization methods are used. The specific mathematical and computational issues introduced for such treatment are given, and a specific case for Computer-Aided Language Learning oriented to the Phonetic Specificities of English for Spanish Speakers is also presented and discussed. This specific technique may also be used in statistically normalizing Speech Data for Speech Recognition Systems.