Identification of ‘‘size-modulated’’ vowels sequences: Effects of modulation periods and speaking rates
Minoru Tsuzaki, Toshio Irino, Roy D. Patterson · The Journal of the Acoustical Society of America · 2005
We investigated the temporal dynamics of auditory normalization and size perception by measuring vowel recognition performance using sequences of vowels in which vocal tract length was modulated during the sequence. The modulation of speaker size was achieved by scaling the frequency axis of the transfer function of vocal tract. The temporal modulation pattern was sinusoidal with a period of 16, 32, 64, 128, 256, 512, 1024, 2048, or 4096 ms. Listeners identified sequences of six vowels from four response alternatives. Although the listeners had no experience with size-modulated speech, the percentage of correct responses was never less than 90% for any modulation period. This suggests that the auditory system has an automatic size-normalizing mechanism which does not require training. The listeners had most difficulty with the 256-ms modulation period, independent of the speaking rate of the sequence. This might indicate a limitation of the processing speed for size-normalization. A simulation using Mellin Images did not reveal any obvious reason for the dip in performance with the 250 ms period, which suggests that the limitation is not in the image construction stage. [Work supported by GASR(A)(2) No. 16200016, JSPS.]