Large Scale Self-Supervised Pretraining for Active Speaker Detection

Otavio Braga, Wei Xia, Keith Johnson, Alice Chuang, Yunfan Ye, Olivier Siohan, Tuấn Anh Nguyễn · 2024

In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show that with pretraining we not only have a more stable supervised training due to better audio-visual features used for initialization, but also improve the ASD mean average precision by 23% on a challenging dataset collected with Google Nest Hub Max devices capturing real user interactions.

Read the paper · More papers on PaperTik