A preliminary study of speech transformation using empirically defined articulatory modes

Brad H. Story, Ingo R. Titze · The Journal of the Acoustical Society of America · 1999

In previous work it has been shown that a speaker-specific set of vocal tract shapes (acquired using MRI) corresponding to vowels can be decomposed into orthogonal components or ‘‘modes,’’ and effectively parameterized by the modal coefficients. Furthermore, a nearly one-to-one mapping was found to exist between the modal coefficients of the two most significant modes and the F1−F2 formant space. Thus an inverse mapping of time-varying formants extracted from recorded speech back to modal coefficients (and consequently to vocal tract area functions) was made possible; the derived sequence of area functions can be used in a simulation of the original speech utterance. In this study, it is first shown that similar orthogonal modes exist for three additional speakers. It will then be demonstrated that the time-varying modal coefficients determined for a recorded sentence via the inverse mapping of one speaker can also be used to create a sequence of area functions based on one of the other speaker’s vocal tracts. A subsequent simulation produces the original sentence but with the vocal tract characteristics of the new speaker. The results also suggest that the time-varying modal coefficients may define common gestures across speakers even though their acoustic characteristics can be different.

Read the paper · More papers on PaperTik