Perceptual evaluation of voicing source models
Jody E. Kreiman, Bruce R. Gerratt, Gang Chen, Mark Garellek, Abeer A. Alwan · The Journal of the Acoustical Society of America · 2012
Many models of the glottal source have been proposed, but none has been systematically validated perceptually, so that it is unclear whether deviations from perfect fit have perceptual importance. If model fit fails in ways that have no perceptual significance, such “errors” can be ignored, but poor fit with respect to perceptually-important features has both theoretical and practical importance. To address this issue, we fit 6 different source models to 40 natural voice sources, and then evaluated fit with respect to time-domain landmarks on the source waveforms and details of the harmonic voice source spectrum. We also generated synthetic copies of the voices using each modeled source pulse, with all other synthesizer parameters held constant, and then conducted a visual sort-and-rate task in which listeners assessed the extent of perceived match between the original natural voice samples and each copy. Discussion will focus on the specific strengths and weaknesses of each modeling approach for characterizing differences in vocal quality. [Work supported by NIH/NIDCD grant DC01797 and NSF grant IIS-1018863.]