Vowel recognition in the absence of formant cues: Dynamic contributions to perception

H. Timothy Bunnell · The Journal of the Acoustical Society of America · 1984

The primary acoustic cue to vowel identity has typically been related to the center frequencies of the first two or three vowel formants. In experiments using specially degraded stimuli, results consistent with this assumption were obtained for stationary synthetic vowels, but not for vowels presented in syllabic contexts (e.g., /VwV/ and /ərVd/). For these experiments, vowel stimuli (both stationary and in context) were synthesized with high-frequency square waves substituted for the standard synthesizer voicing source (Klatt, 1980 cascade synthesizer). Spectra associated with these stimuli have peaks, due to the source harmonics, located at odd multiples of the fundamental frequency. In identification studies using the vowels [ /i/, /ɪ/, /ε/, /æ/, /a/ ], listeners tended to identify stationary vowels on the basis of the frequencies of their largest-magnitude spectral peaks. Since such peaks were invariably coincident with source harmonics, identification errors were frequent. Errors were, however, restricted almost entirely to the vowels [/ɪ/, /ε/, /æ/]. By contrast, vowels in syllabic context were reliably more resistent to identification errors, despite nearly identical acoustic structure throughout the medial portion of the vowel. Results are interpreted to suggest that listeners are well attuned to the contrast between source and vocal tract contributions to the signal. In order to separate these two contributions, however, listeners appeared to need information about changes over time in at least one function (the vocal tract function for these stimuli). [Work supported by NINCDS and the University of Maryland.]

Read the paper · More papers on PaperTik