An Efficient Violence Detection System from Video Clips using ConvLSTM and Keyframe Extraction
Souvik Kumar Parui, Saroj Kr. Biswas, Soumen Das, Manomita Chakraborty, Biswajit Purkayastha · 2023
Monitoring systems are becoming more crucial for city security, specifically to address any kind of violence. Manual detection of violence through video surveillance requires well-trained human resources. However, accuracy and speed are compromised in the case of manual approach. The lack of human resources and their limited ability to interpret violence from a video is the major concern of this approach. Thus, automatic violence detection is required to take timely actions. Most of the existing work for violence detection have used either traditional Machine Learning (ML) based approach or a combination of Convolution Neural Network (CNN) and Long Short Term Memory (LSTM) network for incorporating Spatio-temporal relations. However, video clips contain an extensive number of frames which contain unnecessary and redundant information. Elimination of such frames increases the efficiency and accuracy of the system. Thus, an Efficient ConvLSTM based Violence Detection System (ECLVDS) model is proposed to address this issue. The proposed model uses a clustering-based keyframe extraction technique to eliminate redundant and useless frames. Moreover, the proposed model uses a ConvLSTM model instead of a simple LSTM because the ConvLSTM network extracts better spatiotemporal correlation than the normal LSTM model. The proposed model is evaluated with the widely used benchmark Hockey Fight Dataset and achieved an accuracy of 98.90%, which is comparatively better than some of the existing models.