Video Compression and Action Recognition in Self-supervised Learning
Zongbo Hao, Conghui Hao, Kecheng He · 2024
In computer vision, neural network models typically require a large amount of manually annotated images or video data for training. To reduce annotation costs, self-supervised learning has gained significant attention. This paper proposes a self-supervised learning-based method by introducing an auxiliary task involving spatiotemporal context in videos—extracting video keyframes—to guide self-supervised learning. The neural network learns video stream features by completing this task, thereby understanding the high-level semantics of the video. Additionally, this auxiliary task achieves video content redundancy reduction and video compression through keyframe extraction. To quantitatively evaluate the effectiveness of this method, the neural network is transferred to the domain of action recognition and experiments are conducted on the UCF101 dataset, achieving an accuracy of 66.2% with fewer parameters and training data. The validation of the effectiveness of this method as an auxiliary task for video spatiotemporal features contributes to a better understanding of video semantic features by neural networks and facilitates action recognition, thereby expanding the applications of self-supervised learning.