How (not) to select your voice corpus: random selection vs. phonologically balanced.
Tanya Lambert, Norbert Braunschweiler, Sabine Buchholz · 2007
This paper comparesthe effect of two different voice corpus selection methods on the overall quality of unit selection-based text-to-speech (TTS) voices resulting from training on these corpora. The first selectionmethod aims to maximizethe coverage of stressed as well as unstressed diphones (phonologically balanced: Phonbal) while the second method simply selects sentences at random (Random). We show that, as expected, the Phonbal method results in better phoneticand phonological coverage for the trainingas well as unseen test sentences. However, we also provide evidencefrom an objective evaluationand a subjective listening test that the Random method results in an overall better voice quality when only automatic corpus annotation tools (such as forced alignment)are used, and potentially even with manual annotation. This result has general implications for the fast creation of TTS voices. 1.