Emotional Speech Synthesis

G. Hofer · 2004

The goal of this MSc project was to build unit selection voice that could portray emo-tions in various intensities. A suitable definition of emotion was developed along with a descriptive framework that supported the work carried out. Two speakers were recorded portraying happy and angry speaking styles, additionally a neutral database was also recorded. One voice was built for each speaker that included all the speech from that speaker. A target cost function was implemented that choses units from the database according to emotion mark-up in the database. The Dictionary of Affect [30] supported the emotional target cost function by providing an emotion rating for words in the target utterance. If a word was particularly emotional, units from that emotion were favoured. In addition intensity could be varied which resulted in a bias to select more emotional units. A perceptual evaluation was carried out and subjects were able to recognise reliably, emotions with varying amounts of emotional units present in the target utterance.

Read the paper · More papers on PaperTik