Probabalistic Models and Informativ eS ubspaces for Audiovisual Correspondence

John W. Fisher, Trevor J. Darrell · 2002

We propose a probabalistic model of single source multi- modal generation and sho wh ow algorithms for maximizing mutual infor- mation can find the correspondences between components of each signal. We sho wh ownon-parametric techniques for finding informative sub- spaces can capture the complex statistical relationship between signals in dierent modalities. We extend a previous technique for finding infor- mative subspaces to include ne wp riors on the projection weights, yield- ing more robust results. Applied to human speakers, our model can find the relationship between audio speech and video of facial motion, and partially segment out background events in both channels. We present ne wr esults onthe problem of audio-visual verification, and sho wh ow the audio and video of a speaker can be matched even when no prior model of the speaker's voice or appearance is available.

Read the paper · More papers on PaperTik