Individual variability in the perception of natural and synthetic speech.

Valerie L. Hazan, Bo Shi · The Journal of the Acoustical Society of America · 1992

The extent of individual variability in the perception of natural and synthetic speech was examined in a group of 60 listeners, homogeneous in terms of age, language, background, hearing threshold, and exposure to synthetic speech. Listeners were tested on three types of speech material, representing different levels of contextual information: nonsense syllables (VCV), semantically unpredictable sentences (SUS), and SPIN sentences, in which test words are presented in high- and low-probability sentence contexts. All tests were presented in synthetic speech and natural speech-in-noise conditions. A large amount of variability in scores was found for all tests. Cluster analyses showed the presence of three main listener groups, differing mainly in their performance on low-redundancy sentence material. There was a strong correlation between the SUS and SPIN sentence test results but performance on synthetic speech was not strongly correlated with performance on natural speech in noise. Scores obtained for word and sentence tests were compared to the listeners’ performance on reduced-cue identification tests for place and voicing contrasts in order to ascertain whether there was a link between a listener’s use of acoustic cues and need for contextual redundancy. [Work supported by SERC.]

Read the paper · More papers on PaperTik