TCNet: Gaze Estimation Based on Temporal Body–Head–Eyes Correlation in Dynamic Scenes
Wulue Zhang, Jianbin Xiong, Xiangjun Dong, Qi Wang, Weikun Dai · IEEE Sensors Journal · 2025
In the daily surveillance scenarios, torso, head, and eyes usually cannot be clearly captured by cameras (they may be occluded, blurred, too far away or facing away from the cameras). In the case of occlusion and lacking resolution, most of gaze estimation methods suffer or fall back to approximating gaze with head pose as they excessively rely on clear, close-up views of head and eyes. The method of gaze estimation using full-length images is still under development at recent research. In order to solve these problems and extend this method, we propose a gaze estimation method based on holistic coordination of people. Our method doesn’t depend on clear eyes or face, which is capable of estimating 3-dimensional gaze and providing its confidence for distant views in the case of relying only on a low-pixel body. We associate bodily relationships by encoding the directions and confidences of body, head, gaze into a cascaded Bayesian framework. These directions and confidences are modeled as a von Mises-Fisher distribution. Alongside modeling the spatial coordination of gaze, head and body, we also construct relationships of adjacent frames on the temporal scale. We obtain variable-scale features by 1-dimensional convolution and stochastic depth to adapt to the complicated video inputs. They are utilized to analyze the influence of depth and width on our model. We perform comprehensive experiments of body-head-eyes coordination across various distant views utilizing the GAFA and 3DPW datasets. Finally, we analyze the impact of body and head on gaze based on our testing results.