Multimodal speech synthesis
Juergen Schroeter, Jörn Östermann, Hans Peter Graf, M. Beutnagel, Eric Cosatto, Ann K. Syrdal, Alistair D. Conkie, Y. Stylianon · 2002
Multimodal speech synthesis ("talking heads") encompasses synthesis of speech from text ("text-to-speech", TTS) plus synthesis of a visual presentation of a face that is lip-synced to the generated audio ("visual TTS", VTTS). Talking heads are now practical because of the ever-increasing computing power and falling prices of computer hardware. This paper highlights recent technological breakthroughs relevant to the two modalities. In addition, it exposes synergies between the audio and visual technology components. Finally, the paper summarizes test results that highlight the impact of multimodal speech synthesis in communications and e-commerce applications.