A Unified SpatioTemporal Network with Structural Pruning for Video Action Recognition

Yang-Jie Chen, Rashid Ali, Hsu-Feng Hsiao · 2025

Video action recognition poses significant challenges in capturing and integrating the complex spatiotemporal patterns and motion dynamics necessary for robust understanding. Despite recent advancements, existing deep learning approaches often struggle to efficiently model these interactions over extended temporal ranges. To address this, we propose the Unified SpatioTemporal Network (USTN), a novel framework that fuses segment-level spatiotemporal features with long-range temporal difference information. By strategically employing sparse frame sampling, USTN constructs a rich, coarse-grained representation encapsulating both spatial structure and temporal evolution. Furthermore, we introduce a structural pruning technique to identify and remove redundant parameters, mitigating overfitting and enhancing computational efficiency without compromising performance. Extensive evaluations on the challenging UCF101 and HMDB51 benchmarks, using USTN instantiated with ResNet backbones, demonstrate the superiority of our approach.

Read the paper · More papers on PaperTik