Modeling the perception of vowel-like formant peaks

Hector Raul Javkin, Brian A. Hanson, Paul E. Neyrinck, Hisashi Wakita · The Journal of the Acoustical Society of America · 1987

An improved model of the relationship of harmonic structure to vowel perception is proposed and evaluated. Javkin, Hermansky, and Wakita [11th International Congress of Phonetic Sciences, Tallinn, Estonia, USSR (1987)] compared listeners' responses to one-formant stimuli with the computed weighted average of the two most prominent harmonics [the most important frequency, or MIF of Carlson, Font, and Ganstrom, Auditory, Analysis and Perception of Speech (Academic, London, 1975)] applied with different scaling factors. With this measure, a relatively expanded scale such as magnitude comes closest to perceptual test results. To improve on the characteristics of MIF, a critical band analysis was adopted and Chistovich and Chernova [Speech Common. 5, 3–16 (1986)] were followed in applying a center of gravity formant estimation (CG). Because CG is highly dependent on the integration interval, an iterative process was used, with each iteration choosing the frequency limits for the next analysis until convergence. Initial results suggest that, while CG is less affected by amplitude expansions than MIF, the magnitude space most closely approximates listeners' responses for the one-formant stimuli. Evaluations of the method with multiformant stimuli are also presented.

Read the paper · More papers on PaperTik