Multi-vision encoder with salient foreground separation for video anomaly detection
Suhang Cai, Chong Wang, Sunqi Lin · 2025
Video anomaly detection (VAD) plays a crucial role in intelligent surveillance systems across various public spaces and industries. The visual feature encoder is critical to VAD performance. However, previous works typically employ either a video encoder or an image encoder, limiting their ability to simultaneously recognize both motion and spatial semantics. Additionally, the significant redundancy in surveillance video data remains largely unaddressed. In this paper, we propose a novel multi-vision encoder architecture that combines an image encoder and a video encoder to effectively capture both motion and spatial information. Furthermore, we introduce a salient foreground separation pooling (SFSP) module that directs the model’s attention to motion-significant regions in surveillance videos, enhancing the detection of potentially anomalous events. Extensive experiments demonstrate that our approach achieves competitive performance on two benchmark datasets, UCF-Crime and TAD, validating the effectiveness of our proposed modules.