Speech pattern discovery using audio-visual fusion and canonical correlation analysis
Lei Xie, Yinqing Xu, Lilei Zheng, Qiang Huang, Bingfeng Li · 2012
In this paper, we address the problem of automatic discovery of speech patterns using audio-visual information fusion. Un-like those previous studies based on single audio modality, our work not only uses the acoustic information, but also takes into account the visual features extracted from the mouth region. To improve the effectiveness of the use of multimodal infor-mation, several audio-visual fusion strategies, including feature concatenation, similarity weighting and decision fusion, are uti-lized. Specifically, our decision fusion approach retains the re-liable patterns discovered in the audio and visual modalities. Moreover, we use canonical correlation analysis (CCA) to ad-dress the issue of temporal asynchrony between audio and vi-sual speech modalities and unbounded dynamic time warping (UDTW) is adopted to search for the speech patterns through audio and visual similarity matrices calculated on the aligned audio and visual sequence. Experiments on an audio-visual cor-pus show that, for the first time, speech pattern discovery can be improved by the use of visual information. The decision fusion approach shows superior performance compared with standard feature concatenation and similarity weighting. CCA-based audio-visual synchronization plays an important role in the performance improvement. Index Terms: Speech pattern discovery, canonical correlation analysis, audio-visual speech processing, dynamic time warping