Multi-variety adaptive acoustic modeling in HSMM-based speech synthesis.
Markus Toman, Michael Pucher, Dietmar Schabus · 2013
In this paper we apply adaptive modeling methods in Hid-den Semi-Markov Model (HSMM) based speech synthesis to the modeling of three different varieties, namely standard Aus-trian German, one Middle Bavarian (Upper Austria, Bad Gois-ern), and one South Bavarian (East Tyrol, Innervillgraten) di-alect. We investigate different adaptation methods like dialect-adaptive training and dialect clustering that can exploit the com-mon phone sets of dialects and standard, as well as speaker-dependent modeling. We show that most adaptive and speaker-dependent methods achieve a good score on overall (speaker and variety) similarity. Concerning overall quality there is no significant difference between adaptive methods and speaker-dependent methods in general for the present data set. Index Terms: speech synthesis, dialect, voice modeling, adap-tation