Video-based Driver Action Recognition via Spatial-Temporal and Motion Deep Learning

Fangzhi Ma, Guanyu Xing, Yanli Liu · 2023

Driver action recognition aims to identify different driving actions of drivers, which is of great significance for traffic safety monitoring. However, it is still challenging for the existing methods to identify similar actions which are common situations in driving scenarios. The fundamental reason is that they ignore the global features of videos and fail to extract the most discriminative features. To address this issue, we propose a video-based driver action recognition network via deep learning of spatial-temporal and motion features, which includes two important modules: the global spatial-temporal modeling (GSTM) module and the motion-spatial-temporal joint attention (MSTJA) module. GSTM effectively extracts the global spatial-temporal features of driving actions by expanding the equivalent receptive field of the spatial-temporal dimension through spatial-temporal separable convolution and hierarchical residual connection structures. MSTJA forces the network to focus on the most discriminative areas by jointly exciting the motion patterns of subtle variations and significant spatial-temporal features through the extracted dual path motion features and spatial-temporal features. Experiments demonstrate that the proposed network achieves state-of-the-art classification accuracy on two benchmark driver action datasets Drive&Act and SFD3.

Read the paper · More papers on PaperTik