A vowel feature space from acoustic speech waveforms

H. Wakita · The Journal of the Acoustical Society of America · 1976

Speech sounds can be defined in three major domains: articulatory, acoustic, and perceptual. Some of the parameters in the articulatory domain have already been brought into the acoustic domain for vowel normalization [H. Wakita, J. Acoust. Soc. Am., 57, S3(A) 1975)]. This result can be related to some parameters in the perceptual domain. This paper discusses an attempt to build a unified vowel feature space from acoustic speech waveforms by using currently available techniques. After applying vocal-tract shape and length to normalize the corresponding formant frequencies to those of an acoustic tube having a certain reference length, it is shown that the normalized formant frequencies can be related to the invariant quantities in the perceptual domain which are obtained from multivariate analysis of data in perceptual experiments. Thus, it is possible to define an equivalent vowel feature space in any of the three domains. In its application to automatic acoustic-phonetic transformation for arbitrary speakers, vowel normalization based on vocal tract length can then be performed without directly estimating the length.

Read the paper · More papers on PaperTik