Where Does Gaze Lead? Integrating Gaze and Motion for Enhanced 3D Pose Estimation
Taravat Anvari, Markus Lappe, Marc H. E. de Lussanet · 2025
Understanding and estimating body pose is becoming increasingly important for enhancing user experiences in Virtual Reality (VR). Eye gaze, in particular, plays a critical role in many VR and Augmented Reality (AR) applications. In this paper, we present GaMo, a novel dataset that integrates gaze data and human joint motion capture data during both inter-subject interactions and subject-environment engagements. This dataset provides a comprehensive foundation for advanced pose estimation, enabling the modeling of interactions between users and their surroundings. Based on this dataset, we present the PoseFusionNet model, composed of a Long Short-Term Memory (LSTM) module, and a Transformer Encoder module, focusing on the impact of gaze on body pose estimation. Our model utilizes data from a head-mounted display (HMD), left and right controllers, and 15 previous frames of gaze data to predict the current frame’s pose. Experimental results demonstrate that incorporating gaze data alongside detailed joint information significantly improves pose estimation accuracy.