Utilizing Full-Dimensional Dynamic Convolution for Multi-Modal Gaze Tracking in Conversational Scenarios

Fenglin Huang, Lingfeng Qing · 2024

Gaze tracking estimates the gaze target by interpreting both human behavior and scene information. Existing methods can be categorized into two types: single-mode methods, which rely on scene image analysis but lack temporal context, and multi-mode methods, which combine images, video, and audio cues. While multi-mode methods compensate for the lack of temporal information, they still face limitations in research and accuracy. To improve the accuracy of gaze estimation in conversational scenes, this paper proposes a novel multi-modal gaze tracking method called ADDM. This method integrates full-dimensional dynamic convolution (ODDC) technology, which enhances attention to critical channel information and thereby improves multi-modal feature fusion and tracking accuracy. Additionally, ADDM uses audio as an auxiliary cue to compensate for the lack of temporal information. Comparative and ablation experiments using the VGS datasets show that ADDM increases the AP value by approximately 1.5 with a constant video input size, demonstrating superior accuracy and robustness, thereby proving its effectiveness in complex environments.

Read the paper · More papers on PaperTik