APT: A Simple Adapter for Reusing RGB Video Transformers in Compressed Video Action Recognition
Jiyuan Wang, Huilan Luo · International Journal of Software Engineering and Knowledge Engineering · 2025
The growing adoption of compressed video across diverse applications underscores the demand for efficient action recognition methods. Traditional RGB-based methods face limitations, especially because they depend heavily on computationally intensive optical flow for temporal analysis. We propose a novel method for compressed video action recognition that aims to effectively bridge the gap between compressed-domain and RGB-domain. Specifically, we design a plug-and-play multi-modal feature adapter that enables pretrained RGB-based Transformer models to be directly applied to compressed videos. Our method offers a low-cost cross-domain transfer solution by efficiently fusing I-frame and P-frame within compressed videos, facilitating deep feature alignment and modeling in the compressed domain. Extensive experiments on the UCF101 (93.8%) and HMDB51 (70.7%) datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches for compressed video action recognition.