Speech generation from hand gestures based on space mapping
Aki Kunikoshi, Yu Qiao, Nobuaki Minematsu, Keikichi Hirose · 2009
Individuals with speaking disabilities, particularly people suf-fering from dysarthria, often use a TTS synthesizer for speech communication. Since users always have to type sound symbols and the synthesizer reads them out in a monotonous style, the use of the current synthesizers usually renders real-time opera-tion and lively communication difficult. This is why dysarthric users often fail to control the flow of conversation. In this pa-per, we propose a novel speech generation framework which makes use of hand gestures as input. People usually use tongue gesture transitions for speech generation but we develop a spe-cial glove, by wearing which, speech sounds are generated from hand gesture transitions. For development, GMM-based voice conversion techniques (mapping techniques) are applied to esti-mate a mapping function between a space of hand gestures and another space of speech sounds. In this paper, as an initial trial, a mapping between hand gestures and Japanese vowel sounds is estimated so that topological features of the selected gestures in a feature space and those of the five Japanese vowels in a cepstrum space are equalized. Experiments show that the spe-cial glove can generate good Japanese vowel transitions with voluntary control of duration and articulation. Index Terms: Dysarthria, speech production, hand motions, media conversion, arrangement of gestures and vowels