Dynamic spectral structure really does support vowel recognition
Joanna H. Lowenstein, Susan Nittrouer · The Journal of the Acoustical Society of America · 2011
Traditional approaches to the study of speech perception manipulate cues that are spectrally and temporally discrete, assuming these cues define phonetic segments. Alternatively, dynamic structure across longer signal stretches may support recognition of linguistic units. Strange and colleagues [Strange et al. J. Acoust. Soc. Am. 74, 695–705 (1983)] seemed to support this view by showing that adults can identify vowels in CVCs even when large portions of the syllable centers are missing, but dynamic structure at the margins is preserved. However, adults may just be “filling in the missing parts,” having learned how structure at the margins covaries with structure in the middle. To test these alternatives, adults and 7-yr-olds labeled sine-wave, and vocoded versions of silent-center stimuli. The former preserves dynamic structure; the latter obliterates it. If listeners simply fill in missing parts, adults should perform equally well with sine-wave and vocoded stimuli, and children should perform equally poorly with both. Instead, all listeners performed better with sine-wave stimuli. These outcomes provide further support for the perspective that speech perception is facilitated by broader and longer acoustic structure than that represented by notions of the acoustic cue. [Work supported by NIDCD Grant DC-00633.]