FSConformer: A Frequency-Spatial-Domain CNN-Transformer Two-Stream Network for Compressed Video Action Recognition

Yue Ming, Xiong Gang Lu, Xia Jia, Qingfang Zheng, Jiangwan Zhou · 2023

RGB-based Transformer methods for video action recognition have achieved advanced results recently. However, Transformer lacks local details, leading to the accuracy degradation of small local actions. To alleviate this problem, we propose a frequency-spatial-domain CNN-Transformer two-stream network for compressed video action recognition (FSConformer), which adds frequency-domain local clues to the spatial Transformer. FSConformer takes both frequency and compressed-domain I-frames as input, including a frequency-domain CNN stream and a spatial-domain Transformer stream which capture frequency-domain local features and spatial-domain global features respectively. In the frequency-domain CNN stream, we utilize a frequency-domain spatial-temporal decoder (FDecoder) to integrate multi-scale local features from the frequency-domain CNN backbone and enhance the temporal context. Moreover, we propose a frequency-spatial-domain attentive token fusion (FSAT-Fusion) to combine the complementary frequency-spatial-domain local-global semantics. Experiments on UCF-101, Kinetics-400, and Kinetics-700 reveal that FSConformer reaches higher accuracy compared with other compressed-domain methods. Furthermore, FSConformer achieves competitive accuracy compared with RGB-based Transformer methods in higher inference speed and is superior in small and local actions, which indicates the effectiveness of the local-global complementarity of frequency-domain CNN and spatial-domain Transformer.

Read the paper · More papers on PaperTik