Roles of the average voice in speaker-adaptive HMM-based speech synthesis
Junichi Yamagishi, Oliver Watts, Simon King, Bela Usabaev · 2010
In speaker-adaptive HMM-based speech synthesis, there are a few speakers whose synthetic speech sounds worse than that of other speakers, despite having the same amount of adapta-tion data from within the same corpus. This paper investigates these fluctuations in quality and found that as mel-cepstral dis-tance from the average voice becomes larger, the MOS scores generally become worse. Although the negative correlation ob-tained is not strong enough, this helps us improve the training and adaptation strategies for average voice models. Further-more we remark that this correlation is strongly linked to “vocal attractiveness.” Index Terms: speech synthesis, HMM, average voice, speaker adaptation