Making Mamba Vision Temporal: Leveraging TSM for Efficient Video Understanding
Seung Woo Kwak, Sungjun Hong, Sangyun Lee · 2025
Transformer and Mamba-based models have shown promising results in field of video understanding. However, their excessive complexity often leads to the problem of computational overhead and memory explosion. To address this issue, we propose a novel model called Mamba VisionTSM, which integrates the MambaVision architecture with the Temporal Shift Module (TSM). TSM has been effectively used to enable temporal modeling in 2D CNN models. In this work, we extend its application to the hybrid Mamba Vision model, allowing it to efficiently process temporal information while retaining the long-range spatial dependency capabilities of Mamba Vision. Through experimental results on the Something-Something V2 dataset, we demonstrate that our approach achieves comparable or superior performance compared to existing Transformer-based approaches in action recognition, highlighting its potential for broader video understanding tasks.