Audio-Visual Active Speaker Identification: A comparison of dense image-based features and sparse facial landmark-based features

Warre Geeroms, Gianni Allebosch, Stijn Kindt, Loubna Kadri, Peter Veelaert, Nilesh Madhu · 2022

The field of speaker detection is relatively well researched. Multiple solutions focusing solely on audio or video, or a combination of both exist. On the audio side, a popular feature representation are mel-frequency cepstral coefficients, which are a sparse representation of the audio signal. On the video side, mostly pixel intensities are used, which is not sparse at all. In this paper, we take a look at a sparse video feature representation, namely facial landmarks. We first evaluate what selection of landmarks conveys the most information. Afterwards we propose several neural network architectures trained for audio-visual speaker detection. A comparison both on computational performance and accuracy is shown between the original architecture and architectures utilizing facial landmarks. For the evaluation, we introduce a new dataset to better understand the differences between the pixel and landmark features. The landmark features achieve similar accuracies for a forward oriented head position. There is a small reduction in performance for non-ideal head positions and in the case of occlusions. There is however a significant computational benefit, as there is a complexity reduction of orders two or three of magnitude due to the decreased feature dimensionality. When considering embedded devices, this is a big upside. This way we hope to provide insight and interest in a novel type of active speaker identification models.

Read the paper · More papers on PaperTik