V-APT: Action Recognition Based on Image to Video Prompt-Tuning

Jiangjiang Lan, Tiantian Yuan · 2024

In recent years, the rapid development of 5G network technology and the rise of online short video platforms (e.g., YouTube, TikTok, Kwai) have led to a remarkable increase in the number of online short videos. Action recognition, a fundamental issue in the field of video understanding, has consistently attracted significant attention. The current deep learning action recognition methods, such as C3D and TSN, TimeSformer, etc., are trained to predict a fixed set of predefined categories within a single framework. This pre-defined approach limits the generality and generalization of the model, making it difficult to identify action categories that have not been seen in the video. In this article, we propose an action recognition method based on image-to-video prompt-tuning, namely V-APT. This method renders the CLIP model suitable for video action recognition tasks through prompt fine-tuning. In comparison to existing methods, V-APT requires a minimal number of parameters and exhibits excellent generalization and scalability.

Read the paper · More papers on PaperTik