From high-speed imaging to perception: In search of a perceptually relevant voice source model

Abeer A. Alwan, Patricia Keating, Jody E. Kreiman · The Journal of the Acoustical Society of America · 2011

The goal of this project is to develop a new source model from high-speed recordings of vocal fold vibrations and simultaneous audio recordings. By analyzing area waveforms and acoustics and performing perception experiments, we will better parameterize the new source model and also uncover those aspects of the model that are perceptually salient, leading to improved TTS systems. We recorded acoustic and high-speed video signals from eight speakers producing the vowel /i/. Each speaker produced four phonation types (breathy, modal, pressed, and creaky) and three pitches (low, normal, and high). Glottal area waveforms were extracted from the high-speed recordings, semi-automatically, on a frame-by-frame basis. The acoustic results indicated that phonation types are differentiated by a number of spectral and noise measures, including H1*-H2* and the harmonics-to-noise ratio. In addition, the different pitch levels were responsible for changes in quality within each phonation type, but such changes were subject to inter-speaker variability. Perceptual results highlight the importance of lower-frequency harmonics in voice quality perception. These studies are critical to our computational modeling efforts, because they help us understand which aspects of the source model are perceptually important. [Work supported in part by NSF and NIH.]

Read the paper · More papers on PaperTik