An audio-visual fusion framework with joint dimensionality reducton
Ming Liu, Yun Fu, Thomas S. Huang · IEEE International Conference on Acoustics Speech and Signal Processing · 2008
By combining audio and visual modalities, the speech recognition systems achieve higher performance and robustness. The fusion strategies to this point are mainly three types: feature level fusion model level fusion and decision level fusion. In this paper, we present a novel audio-visual fusion framework, in which a joint dimensionality reductional approach is used to project the audio and visual features into more compact subspaces. With correlation preserving criteria, the representations of projected audio and visual features will be able to preserve the correlation conveyed in the original audio and visual feattre space. At the same time, the better model efficiency is achieved in the more compact feature spaces. The experiments on audio-visual person verification demonstrate the efficiency and effectiveness of the proposed fusion framework.