Expressing vocal effort in concatenative synthesis
Marc L. Schröder, Martine Grice · 2003
A new diphone database with a full diphone set for each of three levels of vocal e#ort is presented. A theoretical motivation is given why this kind of database will be useful for emotional speech synthesis. Two hypotheses are verified in perception experiments: (I) The three diphone sets are perceived as belonging to the same speaker; (II) The vocal e#ort intended during database recordings is perceived in the synthetic voice. The results clearly confirm both hypotheses.