TCAR-Net: Text-Driven Compressed Action Recognition with Multimodal Fusion
Mengkun Guo, Xinqi Li, Die Tao, Ming Ma · Data Intelligence · 2025
Action recognition directly from compressed video streams (leveraging I-frames, motion vectors, residuals) offers significant efficiency gains over pixel-based methods, but faces inherent challenges in achieving deep semantic understanding, especially when integrating rich textual priors from Vision-Language Models (VLMs). The noisy and sparse nature of motion vectors and residuals, complicates direct fusion with fine-grained semantics. To bridge this gap efficiently, this work investigates adapting cross-modal fusion techniques within a parameter-efficient framework tailored for the compressed domain. We introduce TCAR-Net, a text-driven dual-stream architecture built on largely frozen pre-trained backbones. Its spatial branch processes I-frames and residuals, while the motion branch encodes motion vectors (via a frozen ViT); both streams integrate textual guidance using adapted fusion modules. Crucially, only these lightweight crossmodal fusion components are fine-tuned, minimizing adaptation costs. Experiments on UCF101 and HMDB51 demonstrate TCAR-Net achieves competitive accuracy while operating directly on compressed data, significantly reducing decoding overhead. Our findings validate that adapting existing fusion strategies within a parameter-efficient setup is a feasible and effective approach for enabling semantically rich action recognition directly in the compressed domain, offering a practical pathway for efficient video understanding.