Real-Time 3D Gaze Estimation Using Dual-Stream Efficientnet with LSTM and Attention Modules on RGB-D Data

Hafiz Ahmad Qadeer, Minyoung Kim · 2025

Accurate 3D gaze estimation is essential for improving applications in virtual reality, human-computer interaction, and driver monitoring systems. However, existing methods often struggle with challenges such as dynamic lighting, occlusions, and head movements. This study proposes a novel Dual-Stream Gaze Estimation Framework that integrates RGB and Depth (RGB-D) data. The model utilizes EfficientNet-B3 for feature extraction, attention-enhanced fusion for refined feature integration, and a two-layer LSTM for temporal modeling, enabling robust 3D gaze vector prediction (yaw, pitch, and roll). Experiments on the EyeDiap dataset demonstrate that the proposed model achieves a Mean Angular Error (MAE) of 5.96°, representing a 21.2 % improvement over the baseline L2CS-Net (7.56°) and outperforming other state-of-the-art methods. The model maintains real-time performance at 20 FPS. An ablation study confirms the critical role of each component, with the removal of the attention module, depth, and RGB streams increasing MAE to 6.97° 7.02°, and 7.12°, respectively. Furthermore, angular error distribution analysis highlights the model's robustness across diverse conditions. Despite strong performance, the current framework has only been evaluated on the EyeDiap dataset. Future work will focus on optimizing computational efficiency and validating generalizability across more diverse, real-world datasets. Overall, the proposed framework marks a significant advancement in realtime 3D gaze estimation, with promising applications in nextgeneration interactive systems, assistive technologies, and driver monitoring solutions.

Read the paper · More papers on PaperTik