Compressed Video Action Recognition Based on Neural Video Compression

Yuting Mou, Ke Xu, Xinghao Jiang, Tanfeng Sun · 2024

Compressed video action recognition based on traditional codecs, like MPEG-4, H265, etc., has achieved remarkable progress with comparable performance to raw video action recognition. With the development of Neural Video Compression (NVC), action recognition based on NVC should be paid attention to and explored. Firstly, the encoded stream of NVC represents the high-dimension features of the neural network, which allows the features to be utilized for downstream tasks with less additional processing or even directly. Secondly, the high-dimension feature can not be understood by humans, which means the privacy of the raw video frames can be preserved. In this paper, we propose a novel model for compressed video action recognition based on NVC to explore the potential of NVC for action recognition. By introducing spatial and temporal co-attention (ST-CA), the spatial information from the reference frame feature and the temporal information from the motion vector and the residual feature are combined and complemented effectively. The proposed model achieves competitive performance with the traditional compressed video action recognition methods and the raw video action recognition methods on the HMDB-51 and UCF-101 datasets. Besides, the proposed model preserves the privacy of the raw video frames and has much less computational complexity than the raw video action recognition methods.

Read the paper · More papers on PaperTik