FusGaze: Full range gaze estimation with multi-scale fusion

Dayeon Yoo, Jaehun Cho, Kwangho Song · Computer Vision and Image Understanding · 2026

Gaze estimation plays a crucial role in a wide range of applications, including human–computer interaction (HCI), augmented and virtual reality (AR/VR), driver monitoring and robotic intelligence. Most of the conventional approaches rely on external detectors to extract head or face regions prior to gaze estimation. It causes a fundamental issue by inheriting the detector errors, which leads to incorrect gaze predictions based on inaccurate input data. To overcome these limitations, we propose FusGaze, a novel end-to-end gaze estimation framework that seamlessly integrates detection and estimation within a single architecture. FusGaze eliminates external detector dependencies by appending an object detection head and a gaze estimation head to a single detection backbone, thereby enabling more stable and robust gaze estimation in practical environments. To enable effective multi-scale fusion, we design an Adaptive Weighted Late Fusion (AWLF) block that data-drivenly fuses multi-scale gaze vectors. Moreover, a 2D Gaussian heatmap-based auxiliary loss function is introduced to improve estimation stability, even under challenging scenarios such as side or rear views where visual cues are limited. Experiments conducted on the Gaze360 benchmark demonstrate that FusGaze achieves state-of-the-art performance with mean angular error of 12.85°, among single image-based gaze estimation models. In particular, the model shows significant improvements under front-facing conditions while maintaining competitive performance under backward-facing conditions, thus experimentally validating its effectiveness in full 360° gaze estimation.

Read the paper · More papers on PaperTik