Multi-Stream Feature Fusion Framework with Separable Convolutional LSTM for Violence Detection in Surveillance Video

B Sushma, Dinesh Kumar, M. Murali, T. V. Sushma · 2025

Violence detection in video surveillance is a critical task for ensuring public safety and effective monitoring. Traditional methods often struggle with accurately identifying violent actions due to the complexity of capturing both spatial, motion-based and temporal features, leading to suboptimal performance in dynamic environments. The proposed work introduces a multi-stream feature fusion framework that integrates spatial, temporal, and motion-based features to achieve robust and efficient violence detection. The spatial and appearance features are captured from the RGB stream, while the depth stream adds valuable 3D context for object positioning and proximity analysis. The optical flow stream focuses on motion dynamics to highlight rapid and violent movements while the background subtraction stream isolates foreground movements by reducing background noise for more accurate detection of dynamic actions. The features extracted from each stream are temporal modelled by Separable Convolutional LSTM (S-ConvLSTM) to effectively capture the spatio-temporal dynamics. The spatio-temporal features are effectively fused to improve the violence and non-violence classification accuracy and performance.

Read the paper · More papers on PaperTik