Evaluation of the expressivity of a Swedish talking head in the context of human-machine interaction

Jonas Beskow, Loredana Cerrato · 2008

Una delle piu ’ recenti sfide nell´ambito dello sviluppo di sistemi automatici per la riproduzione di parlato audio-visivo è quella di riuscire a sviluppare un modello per la produzione di parlato espressivo bi-modale. In questa comunicazione verranno presentati i risultati del primo tentativo di far produrre ad una testa parlante svedese una sintesi audiovisiva di parlato espressivo e si discuterá della messa a punto e dei risultati di due test percettivi condotti allo scopo di valutare l´espressivitá di questa testa parlante, inserita in un contesto simulato di interazione uomo-macchina. This paper describes a first attempt at synthesis and evaluation of expressive visual articulation using an MPEG-4 based virtual talking head. The synthesis is data-driven, trained on a corpus of emotional speech recorded using optical motion capture. Each emotion is modelled separately using principal component analysis and a parametric coarticulation model. In order to evaluate the expressivity of the data driven synthesis two tests were conducted. Our talking head was used in interactions with a human being in a given realistic usage context. The interactions were presented to external observers that were asked to judge the emotion of the talking head. The participants in the experiment could only hear the voice of the user, which was a pre-recorded female voice, and see and hear the talking head. The results of the evaluation, even if constrained by the results of the implementation, clearly show that the visual expression plays a relevant role in the recognition of emotions.

Read the paper · More papers on PaperTik