Novel Applications of Neural Networks in Speech Technology Systems: Search Space Reduction and Prosodic Modeling
Javier Macías-Guarasa, Juan Manuel Montero, Javier Ferreiros, Ricardo de Córdoba, Rubén San-Segundo, Juana María Gutiérrez-Arriola, Luis Fernando D’Haro, Fernando Fernández-Martínez, Roberto Barra-Chicote, José Manuel Pardo · Intelligent Automation & Soft Computing · 2009
Abstract Neural networks (NNs) have been extensively used in speech technology systems. In this paper, we present two novel applications of NNs in speech recognition and text-to-speech systems. The prosodic modeling is one of the most important tasks for developing a new text-tospeech synthesizer, especially in a female-voice high-quality restricted-domain system. Our double objective is to get accurate predictors for both the fundamental frequency (FO) curve and phoneme duration by minimizing the model estimation error in a Spanish text-to-speech system, by means of a neural network estimator, which has proved to be an excellent tool for the modeling. The resulting system predicts prosody with very good results (for duration: 15.5 ms in RMS and a correlation factor of 0.8975; for F0: 19.80 Hz in RMS and a relative RMS error of 0.43) that clearly improves our previous rule-based system.