From audio-only to audio and video text-to-speech
Eric Cosatto, Hans Peter Graf, Jörn Östermann, Juergen Schroeter · 2004
Assessing the quality of Text-to-Speech (TTS) systems is a complex problem due to the many modules involved that address different subtasks during synthesis. Adding face synthesis – the animation of a “talking head ” and its rendering to video – to a TTS system makes evaluation even more difficult. In the case of talking heads, today, we are at the infancy of research towards evaluating such systems. This paper reports on progress made with the AT&T sample-based Visual TTS (VTTS) system. Our system incorporates unit-selection synthesis (now well known from Audio TTS) and a moderate-size recorded database of video segments that are modified and concatenated to render the desired output. Given the high quality the system achieves, we feel for the first time that we are close to passing the Turing test, that is, that we are almost able to synthesize “talking heads ” that look like recordings of real people. We demonstrate this point in applications, either over the web (client/server), or in stand-alone form, in a kiosk setting. Several steps are necessary to assure a very high quality sample based VTTS system. First, highly accurate image analysis tools are important for creating the necessary video clip databases. The problem is compounded by the fact that facial videos cannot be stored whole due to unfavorable combinatorics: for a given synthetic sequence, it is very unlikely that a whole face video clip contains the correct mouth sequence, the appropriate eye sequence, and also a suitable “background ” face, given what we want to synthesize. Consequently, separate parts of a synthetic face need to be accessible independently from each other