On invariance: Acoustic input meets listener expectations
Allard Jongman, Bob McMurray · 2017
Speech perception has been classically framed in terms of the widespread variability in speech acoustics. Factors like speaking rate, coarticulation, and speaker affect virtually all phonetic measurements or “cues”. However, our understanding of this problem has been built on the basis of small-scale phonetic work, one cue and context at a time. We present findings based on a large corpus of fricatives that suggest that the massive variability in speech may not be insurmountable, but rather can be described as the simple additive product of multiple known factors. At any given moment, listeners have expectations about the anticipated value of cues like formant frequency or fricative spectrum as a function of contextual factors like talker and vowel. Perception is then based on the difference between the actual cue values heard and these expectations. We briefly describe two additional experiments that demonstrate that manipulation of listeners’ expectations can change the accuracy of fricative identification, and improve listeners’ ability to predict the subsequent vowel. Finally, we present new data on the relative contributions of place and voicing cues which suggest that there are several acoustic cues that can be considered invariant. However, this information alone is not sufficient to account for listeners’ identification of fricatives. To approximate the performance of human listeners requires many cues, and these cues need to be interpreted relative to expectations derived from context.