Temporal characteristics of extraction of size information in speech sounds

Chihiro Takeshima, Minoru Tsuzaki, Toshio Irino · The Journal of the Acoustical Society of America · 2006

We can identify vowels pronounced by speakers with any size vocal tract. Together, we can discriminate the different sizes of vocal tracts. To simulate these abilities, a computational model has been proposed in which size information is extracted and separated from the shape information. It is important to investigate temporal characteristics of the size extraction process. Experiments were performed for listeners to detect the size modulation in vowel sequences. All the sequences had six segments. Each segment contained one of three Japanese vowels: ‘‘a,’’ ‘‘i,’’ and ‘‘u.’’ Size modulation was applied by dilating or compressing the frequency axis of continuous, STRAIGHT spectra. Modulation was achieved by changing the dilation/compression factor in sinusoidal functions. The original F0 pattern of the base sequence, except for warping of the time axis, was used for all stimuli. The minimum modulation depth at which listeners were able to detect the existence of modulation was measured as a function of the modulation frequency. The results will be compared with low-pass characteristics in a temporal modulation transfer function obtained with the amplitude-modulated noise. They will be discussed in relation to a computational model based on the Mellin transformation.

Read the paper · More papers on PaperTik