EgoNet: An Unified Egocentric Active Speaker Detection Framework for both Camera Wearer and Visible Candidates

Yongqian Li, Xin Zhou, Zheng He, Wei Yu, Yong Huan Luo · 2025

Active Speaker Detection (ASD) aims to determine whether each candidate in a video frame is speaking. The egocentric dataset Ego4D introduces unique challenges for this task, such as dynamic shooting angles that cause candidates to frequently leave the sight, leading to temporal discontinuities. Additionally, Ego4D poses a novel task: detecting the speaking activities of the camera wearer, who never appears in the field of view. Existing methods treat these two tasks separately, and treat candidates out of sight as noise. In contrast, we propose EgoNet, a framework that uniformly models all candidates, including those not visible. By capturing interactions among all candidates and modeling broader temporal context, EgoNet reduces uncertainty and improves performance in egocentric active speaker detection.

Read the paper · More papers on PaperTik