Speech estimation in non-stationary noise environments using timing structures between mouth movements and sound signals
Hiroaki Kawashima, Yu Horii, Takashi Matsuyama · 2010
A variety of methods for audio-visual integration, which inte-grate audio and visual information at the level of either features, states, or classifier outputs, have been proposed for the purpose of robust speech recognition. However, these methods do not al-ways fully utilize auditory information when the signal-to-noise ratio becomes low. In this paper, we propose a novel approach to estimate speech signal in noise environments. The key idea behind this approach is to exploit clean speech candidates gen-erated by using timing structures between mouth movements and sound signals. We first extract a pair of feature sequences of media signals and segment each sequence into temporal inter-vals. Then, we construct a cross-media timing-structure model of human speech by learning the temporal relations of overlap-ping intervals. Based on the learned model, we generate clean speech candidates from the observed mouth movements. Index Terms: multimodal, non-stationary noise, timing, linear dynamical system, particle filtering