Fusing audio and video information for online speaker diarization
Joerg Schmalenstroeer, Martin Kelling, Volker Leutnant, Reinhold Haeb‐Umbach · 2009
In this paper we present a system for identifying and localizing speakers using distant microphone arrays and a steerable pantilt-zoom camera. The scenario at hand assumes audio streams to be processed in real-time to get the diarization information “who spokes when and where ” with only short delays. Our new idea is to fuse the acoustical and visual observations directly within the Viterbi decoder to improve the diarization process. In contrast to standard Viterbi decoder implementations, we use a time variant transition matrix generated from speaker change hypotheses and location information. This allows a simultaneous segmentation and classification of the audio stream. Experiments show, that video information enables a substantial improvement of the diarization results. Index Terms: speaker diarization, face identification, acoustic scene analysis