Audio-Driven Gaze Estimation for MOOC Video
Bowen Tian, Long Rao, Wei Xu, Wenqing Cheng · 2023
Gaze estimation, which refers to predicting the gaze direction of human eyes, is essential in users state awareness and human attention estimation. Most existing methods only use the images of eyes or faces to obtain a good result, which is not enough in some teaching scenarios. In order to be more in line with the teaching scenarios, some additional information could have been used to correct the gaze model predictions, including the video and audio content viewed by the user. Therefore, we propose a new dataset, denoted as MOOCGaze dataset, which includes time-synchronized screen recordings, the corresponding audio, user-facing camera views, and eye gaze data. We also select a state-of-the-art model to obtain the initial Point-of-Gaze and introduce the RefineNet with an audio-video modality fusion scheme to correct the prediction results. Our final method yields significant improvements on our MOOCGaze dataset, with 26.1% improvement in the prediction of eye direction (resulting in 1.947 degrees in angular error).