Enhancing CCTV Violence Detection: A Comparative Study of Deep Learning Models for Violence Detection in Surveillance Videos

K. Maxim Rodrigues, Glen Dsouza, Omkar Phansopkar, Pramod J. Bide · 2024

With the large number of CCTV cameras located worldwide, ensuring people's safety has become much easier. Despite this, it is impossible to keep track of 100s of CCTV cameras simultaneously. Therefore, deep learning methods of violence detection from surveillance footage have been proposed to reduce human intervention. This solution presents a comparison of the performance of three-dimensional (3D) Convolutional Neural Networks (CNNs), VGG networks, and MobileNet models combined with Long Short-Term Memory (LSTM) units and Multi-Stream Networks to detect violence in CCTV video footage, providing a comprehensive analysis of these models based on accuracy, hardware requirements, and other performance metrics essential to real-world applications. The majority of the existing systems are trained on violent footage from sports or movies, the models in this solution are trained on a combination of datasets consisting of CCTV footage which provides an accurate representation of real-world surveillance footage, thus improving model accuracy. Existing research does not study the effect of combining very deep CNN models like VGG with advanced techniques like multi-stream models. This solution explores the impact of integrating such models to create a more robust solution. All the models proposed and evaluated have produced accurate results and are capable of identifying violence based on surveillance footage, with an accuracy of up to 98.46%. Unlike other solutions that do not take into account the memory and prediction time of different models, this solution identifies the ideal model depending on the available resources and infrastructure, helping make a decision based on real-world requirements.

Read the paper · More papers on PaperTik