Multi-Speaker Tracking by Fusing Audio and Video Information

Zichao Xiong, Hongqing Liu, Yi Zhou, Zhen Luo · 2021

In this work, a novel approach is proposed for multi-speaker tracking by integrating audio and visual data in a particle filtering framework. To that aim, first, the audio signal is measured based on the maximum Global Coherence Field (GCF) peak at each frame, and the visual signal is measured using a face detector. Second, we exploit visual observations (object detections) to guide the acoustic processing by constraining the acoustic likelihood on the horizontal plane defined by the predicted height of a speaker. Moreover, the video likelihood function is improved by fusing audio positioning information and face detection 3D mouth information. In particular, we set an adjustable weight to better integrate audio positioning information and 3D mouth information in different scenes. A particle filtering framework and greedy data association are used to incorporate visual and acoustic likelihoods to track multi-speakers. Finally, the CAV3D dataset is used to demonstrate the performance of the proposed method in a variety of scenarios such as occlusion and outside the Field of View (FoV).

Read the paper · More papers on PaperTik