Joint processing of audio and visual information for multimedia indexing and human-computer interaction
C. Neti, Benoit Maison, Andrew Senior, Giridharan Iyengar, P. Decuetos, Srabanti Basu, Ashish Kumar Verma · 2000
Information fusion in the context of combining multiple streams of data e.g., audio streams and video streams corresponding to the same perceptual process is considered in a somewhat generalized setting. Specifically, we consider the problem of combining visual cues with audio signals for the purpose of improved automatic machine recognition of descriptors e.g., speech recognition/transcription, speaker change detection, speaker identification and speaker event detection. These happen to be important descriptors for multimedia content (video) for efficient search and retrieval. A general framework for considering all of these fusion problems in a unified setting is considered.