Seeing Speech: Capturing Vocal Tract Shaping Using Real-Time Magnetic Resonance Imaging

Erik Bresch, Yoon‐Chul Kim, Krishna S. Nayak, Dani Byrd, Shrikanth S. Narayanan · 2008

Understanding human speech production is of great interest from engineering, linguistic, and several other research points of view. While several types of data available to speech understanding studies lead to different avenues for research, in this article we focus on real-time (RT) magnetic resonance imaging (MRI) as an emerging technique for studying speech production. We discuss the details and challenges of RT magnetic resonance (MR) acquisition and analysis, and modeling approaches that make use of MRI data for studying speech production. MOTIVATION From an engineer’s point of view, detailed knowledge about speech production gives rise to refined models for the speech signal that can be exploited for the design of powerful speech recognition, coding, and synthesis systems. From a linguist’s point of view, speech research may be conducted to address open questions in the areas of phonetics and phonology. These include: 1) what articulatory mechanisms explain the inter- and intrasubject variability of speech, 2) what aspects of the vocal tract shaping are critically controlled by the brain for conveying meaning and emotions, and 3) how does prosody affect the articulatory timing. From other research points of view, speech production is important to understand language acquisition and language disorders. All of these efforts require intimate knowledge of the speech generation mechanisms. Different types of data are available to the speech researcher—from audio and video recordings of speech production to muscle activity data produced by electromyography, respiratory data from subglottal or interoral pressure transduction, and images of the larynx obtained through video laryngoscopy. While the vocal tract posture and movement can be investigated using a host of techniques summarized in Table 1 including X ray (microbeam), cinefluography, ultrasound, palatography, electromagnetometry (EMA), RT-MRI has a particular advantage in that it produces complete views of the entire vocal tract including the pharyngeal structures in a safe and noninvasive manner. With RT-MRI, a midsaggital image of the vocal tract from the glottis (bottom) to the lips (left) can be acquired as illustrated in Figure 1(a). In this image, we can trace the air-tissue boundaries of the anatomical components that are of interest to the speech researcher and obtain a representation similar to Figure 1(b). These components, also known as articulators, are controlled by the brain during speech production and are used to change the shape of the vocal tract tube. With it, they also change the filter function for the excitation signal generated at the glottis and elsewhere along the airway. Hence the motion of the articulators shapes the sounds of speech and other human vocalizations. The signal processing challenges when studying these using RT-MRI lie in the fast acquisition of high-quality RT MRI images including simultaneous noise-robust audio recording [1], the subsequent detection of the relevant features from each image, and the analysis and modeling of the time-varying vocal tract shape for the purpose of gaining deeper understanding of the underlying principles that govern the speech production process.

Read the paper · More papers on PaperTik