State Space Model Based VideoMAE Enhancement for Efficient Video Action Classification
Junbeom Moon, Sehwan Heo, Jiye Won, Jaeseok Jang, Soon Ki Jung · 2025
Video action classification is a critical task in computer vision, with applications spanning security surveillance, sports analytics, and human-computer interaction. Recent advancements, such as VideoMAE, have demonstrated the effectiveness of transformer-based architectures in learning spatiotemporal representations. However, their computational inefficiency, due to quadratic time complexity, limits their applicability in resource-constrained environments. This study addresses these limitations by introducing Vision Mamba, an efficient encoder based on bidirectional state-space models (SSMs), into the VideoMAE framework. Vision Mamba processes sequences in both forward and backward directions, enabling robust temporal modeling with linear complexity. Experimental results show that Vision Mamba reduces memory consumption by up to 48.9% compared to traditional transformer-based methods, while achieving a 1.59% improvement in Top 1 and 2.01% improvement in Top5 classification accuracy on the Kinetics-400 dataset. These findings underscore the potential of Vision Mamba as a scalable and resource-efficient solution for video action classification.