Temporal Shift Vision Transformer Adapter for Efficient Video Action Recognition
Yaning Shi, Pu Sun, Bing Gu, Longfei Li · 2024
In recent work, the performance of video action recognition based on Vision Transformer (ViT) far exceeds that of CNN-based methods. Training the video ViT model requires larger datasets and more GPU computing resources. How to reduce ViT's dependence on large-scale datasets and computing resources has become a widespread concern among researchers. We propose a Temporal Shift Adapter (TSM-Adapter) to efficiently transfer large-scale pre-trained ViT parameters to downstream video action recognition tasks. Through experiments on the smaller video behavior recognition dataset TinySthv2, compared with the method of completely fine-tuning ViT, our TSM-Adapter only trained 16.5% of the newly introduced model parameters and achieved a Top-1 accuracy improvement of 12.87%. Moreover, compared with similar methods, under the same training parameters, our method achieved a 1.4% improvement in Top-1 recognition accuracy.