Domain adaptation methods in the IBM trainable text-to-speech system
Volker Fischer, Jaime Botella Ordinas, Siegfried Kunzmann · 2004
This paper presents a comparison of domain adaptation techniques for a unit selection based text-to-speech system. The methods under investigation consider two different pre-requisites, namely the absence and the existence of addi-tional domain specific training prompts, spoken by the orig-inal voice talent. Whereas in the first case we employ do-main specific pre-selection, for the latter we compare a va-riety of methods that range from a simple extension of the segment inventory to a complete reconstruction of the sys-tem, which also includes the training of decision trees for the domain dependent prediction of prosody targets. An ex-perimental evaluation of the methods under consideration unveils significant improvements (up to 1.1 on a 5 point MOS scale) over the baseline system for sentences from the target domain, while showing no significant degradation when synthesizing sentences from other than the adaptation domain. 1.