Should ASR front-end be insensitive to fundamental frequency? (perceptual shift of formant position due to fine harmonic structure of voiced speech

Hynek Heřmanský · The Journal of the Acoustical Society of America · 1987

Due to the fine harmonic structure of the source spectrum of voiced speech, most automatic speech analysis techniques estimate formant positions that systematically deviate from the resonant frequencies of the vocal tract. This property is in conflict with linear source-filter models of speech production and an effort is made to minimize the effect of the fundamental frequency F0 on the analysis result [e.g., Hermansky et al., Proceedings ICASSP-83 (IEEE, New York, 1983), pp. 778–780]. Our experiments show that the human auditory perception has properties similar to the above described properties of automatic analyses. The method of confusion scaling is applied in paired comparison of synthetic vowel-like stimuli with varying F0 and the second formant frequency F2. The results show that the perceptual estimate of formant position goes through a cycle of positive and negative deviations as the relative F2/F0 frequency goes from one integral number to the next. The position is estimated the most accurately when the formant peak falls either into a harmonic peak or if it lies right between two harmonic peaks. The maximal deviation of the estimate due to harmonics of the F0 from the 120-Hz neighborhood in the 2-kHz second formant region is about ± 0.5%. Our findings imply that when modeling the speech perception, as in the automatic speech recognition, a certain degree of sensitivity of the analysis technique to F0 might be beneficial.

Read the paper · More papers on PaperTik