Identification of synthetic vowels based on time-varying vocal tract area functions.

Kate Bunton, Brad H. Story · The Journal of the Acoustical Society of America · 2009

Identification accuracy for synthetic vowel stimuli generated with static vocal tract area functions from eight speakers was recently reported (K. Bunton and B. H. Story, JASA, in press). Although vowels were identified with reasonably high accuracy, in many cases neighboring vowels were confused. In the present study, new stimuli were generated based, again, on area functions from the same eight speakers. This time, however, the temporal variation of the vocal tract shape during natural production of isolated vowels was incorporated by developing a formant-to-area mapping for each speaker based on principal components analysis. This allowed formant contours extracted from recorded vowels to be mapped onto time-varying area functions. These were coupled to a voice source model and acoustic waves were propagated with a wave-reflection vocal tract model to generate vowels. Vowels were identified by listeners using a forced choice paradigm. Results indicated that including variation in vowel tract shapes, and hence formant frequencies, along with natural vowel durations improved identification accuracy relative to the previously tested static versions. It is concluded that these eight sets of area functions generate good renditions of target vowels when changes in vocal tract shape and duration were included. [Research supported by NIH R01-DC04789.]

Read the paper · More papers on PaperTik