Driving Video Fixation Prediction Model Via Spatio-Temporal Networks and Attention Gates
Tao Deng, Fei Yan, Hongmei Yan · 2021
Driving fixation prediction is becoming an essential research problem in human-like driving systems or advanced driver assistance systems (ADAS) in a dynamic driving environment. However, it is still a lack of driving video fixation prediction models that can dynamically predict drivers’ fixational locations. In this work, we propose a driving video fixation prediction model via spatio-temporal networks and attention gates method, named as DSTANet, to predict drivers’ attention in the dynamic driving videos. The spatial and temporal driving information are both considered in DSTANet by convolutional long short-term memory (ConvLSTM). In addition, the human attention mechanism is designed to filter some driving-irrelevant information via the attention gates (AGs). The experimental results indicate that the proposed DSTANet outperforms the state-of-the-art saliency models and predicts drivers’ attentional spatial locations more accurately. Furthermore, the time-series prediction results show that DSTANet is more robust and coherent than others and includes more temporal information.