Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech
Gustav Eje Henter, Thomas Wayne Merritt, Matt Shannon, Catherine Mayo, Simon King · 2014
Acoustic models used for statistical parametric speech synthe-sis typically incorporate many modelling assumptions. It is an open question to what extent these assumptions limit the natu-ralness of synthesised speech. To investigate this question, we recorded a speech corpus where each prompt was read aloud multiple times. By combining speech parameter trajectories ex-tracted from different repetitions, we were able to quantify the perceptual effects of certain commonly used modelling assump-tions. Subjective listening tests show that taking the source and filter parameters to be conditionally independent, or using di-agonal covariance matrices, significantly limits the naturalness that can be achieved. Our experimental results also demonstrate the shortcomings of mean-based parameter generation. Index terms: speech synthesis, acoustic modelling, stream in-dependence, diagonal covariance matrices, repeated speech 1.