Audio-visual anticipatory coarticulation modeling by human and machine

Louis H. Terry, Karen Livescu, Janet B. Pierrehumbert, Aggelos K. Katsaggelos · 2010

The phenomenon of anticipatory coarticulation provides a ba-sis for the observed asynchrony between the acoustic and vi-sual onsets of phones in certain linguistic contexts. This type of asynchrony is typically not explicitly modeled in audio-visual speech models. In this work, we study within-word audio-visual asynchrony using manual labels of words in which theory suggests that audio-visual asynchrony should occur, and show that these hand labels confirm the theory. We then introduce a new statistical model of audio-visual speech, the asynchrony-dependent transition (ADT) model. This model allows asyn-chrony between audio and video states within word boundaries, where the audio and video state transitions depend not only on the state of that modality, but also on the instantaneous asyn-chrony. The ADT model outperforms a baseline synchronous model in mimicking the hand labels in a forced alignment task, and its behavior as parameters are changed conforms to our ex-pectations about anticipatory coarticulation. The same model could be used for speech recognition, although here we consider it only for the task of forced alignment for linguistic analysis. Index Terms: audio-visual speech recognition, audio-visual asynchrony, anticipatory coarticulation, dynamic Bayesian net-works 1.

Read the paper · More papers on PaperTik