Optimal event search using a structural cost function - improvement of structure to speech conversion
Daisuke Saito, Yu Qiao, Nobuaki Minematsu, Keikichi Hirose · 2009
This paper describes a new and improved method for the frame-work of structure to speech conversion we previously proposed. Most of the speech synthesizers take a phoneme sequence as input and generate speech by converting each of the phonemes into its corresponding sound. In other words, they simulate a human process of reading text out. However, infants usually acquire speech communication ability without text or phoneme sequences. Since their phonemic awareness is very immature, they can hardly decompose an utterance into a sequence of phones or phonemes. As developmental psychology claims, in-fants acquire the holistic sound patterns of words from the utter-ances of their parents, called word Gestalt, and they reproduce them with their vocal tubes. This behavior is called vocal im-itation. In our previous studies, the word Gestalt was defined physically and a method of extracting it from a word utterance was proposed. We already applied the word Gestalt to ASR, CALL, and also speech generation, which we call structure to speech conversion. Unlike reading machines, our framework simulates infants ’ vocal imitation. In this paper, a method for improving our speech generation framework based on a struc-tural cost function is proposed and evaluated. Index Terms: speech synthesis, the structural representation, vocal imitation, a structural cost function