The unwarranted effectiveness of smoothness as a criterion for estimating vocal-tract length from speech
Edward P. Neuburg · The Journal of the Acoustical Society of America · 2001
Unambiguous quantitative characterization of speech sounds would be a useful thing. This talk will be about experiments done in an attempt to find such a characterization, by studying corpora of spoken vowel tokens all of which ‘‘sound alike.’’ Using Linear Predictive Coding (LPC) of order 10, 4 damped sinusoids (‘‘formants’’) were recovered from the tokens, and their bandwidths were algorithmically fixed as a function of frequency. (Other formant finders have also been tried.) Hypothesizing that a token is composed of 4 damped sinusoids is equivalent to hypothesizing that it was produced by a ‘‘vocal tract’’ with 8 sections; the shape (8 areas) of that tract is a function of its length, which is a free parameter of the formant-to-tract transformation. The experiments show that over a corpus of tokens of a given speech sound, simply choosing the vocal-tract length that produces the ‘‘smoothest’’ tract will make for shapes that are very nearly identical over all tokens, for all talkers (in particular for men and women), and are physiologically reasonable. Such 8-parameter shapes might be candidates for the desired characterization.