Localization of text-to-speech voices in a virtual speech display
Gordon Rubin · The Journal of the Acoustical Society of America · 2006
An auditory display using high-quality text-to-speech (TTS) voices is presented to a group of individuals. Subjects are asked for their impression of the source direction of a synthesized talker in a nonreverberant environment. Location is coded using a set of nonindividualized head-related transfer functions (HRTFs). Many years of research in binaural technology provide comparative data for the localization ambiguities inherent in such presentation (e.g., front-back confusion). These ambiguities are characterized for stimuli generated by a mature concatenative speech synthesis system. An initial attempt is made to quantify the perceptual significance realized by a binaural-model-motivated speech synthesizer.