Toward data-driven modeling of dynamic vocal-tract data

Abhinav Sethy, Shrikanth Shri Narayanan, Sungbok Lee, Dani Byrd · The Journal of the Acoustical Society of America · 2003

Speech production modeling relies on various forms of biometric measurements for getting articulatory data. Recent advances in real-time magnetic resonance imaging (rtMRI) capabilities [Narayanan et al., J. Acoust. Soc. Am. 113, 2258 (2003)] which make it possible to capture vocal tract images at 24-fps promise to illuminate finer details of speech production. This study presents our first effort toward automatically constructing statistical models of the vocal tract from rtMRI data. We present an automated method of extracting the regions of interest from image sequences based on Kalman snakes and optical flow [Cootes et al., IVCV July 1994]. The time resolution of this image tracking system is improved by incorporating information from parallel vocal tract movement data obtained by the electromagnetic articulography (EMA) system [Perkell et al., J. Acoust. Soc. Am. 92]. EMA data provide kinematic information such as velocity and acceleration at several points along the vocal tract. Data from segmentation and tracking are merged with information from EMA and from acoustic analysis of speech in order to obtain a time series of feature vectors. Feature space reduction and statistical parametrization schemes such as PCA are then applied to get parametric models for the articulatory system. [Work supported by NIH Grant No. DC 03172.]

Read the paper · More papers on PaperTik