STDN: A SpatioTemporal Difference Network for Video Action Recognition

Zhao Guo, Yuelei Xiao, yi Jun Li, Cheng Fan · 2023

Video action recognition is a challenging task in the realm of computer vision, which necessitates the acquisition of diverse and dynamic temporal and spatial information. Due to the promising results in capturing temporal dynamics, the Temporal Difference Network (TDN) has become an important method for video action recognition. However, it still encounters issues in effectively fusing spatiotemporal information. To overcome these issues, this study proposes an innovative SpatioTemporal Difference Network (STDN) for video action recognition based on spatiotemporal information fusion. The proposed STDN method improves upon TDN by devising and integrating spatiotemporal residual block (MSRB), each comprising a Multi-path Excitation (MPE) module. The MPE module is tailored to selectively activate informative features and suppress unimportant ones through adaptive feature weighting. Thus, it enhances the network's discriminative power by capturing more diverse spatiotemporal feature information and modeling inter-channel relationships. The experimental results demonstrate the effectiveness of our proposed STDN method on standard benchmarks, achieving 2.9% and 1.2% improvement on the Something-Something V1 & V2 datasets, respectively.

Read the paper · More papers on PaperTik