A Vision Transformer Model for Violence Detection from Real-Time Videos
Arfin Shagufta, Mohammad Tarique Hesham, Sarfaraz Masood, Ahmed A. Abd El‐Latif · The 5th International Conference on Future Networks & Distributed Systems · 2021
Based on the rising incidences of crime and violence, it has become a matter of general importance that technology may be developed to automatically detect the presence of violence in the surveillance footage. For law enforcement, the detection of violent incidents can play an important role in urban safety. The efficiency of event detectors is usually measured in terms of detection speed, precision, and generality over many types of video inputs in a distinct format. However, various recent studies in this area have either focused on the correctness of the model, its response speed, or both, but have missed on considering the multiple data sources for effective model analysis. The main objective of the work is to propose a real-time violence event identifier based on state-of-the-art deep learning methods. Primarily the proposed model is based on the ViT (Vision Transformer) architecture, while other models like ConvLSTM and VGG16 with LSTM were also explored in this work. Two benchmark datasets viz. the RLVS (Real Life Violence Situations) and the Hockey fight datasets were used in this work for robust training and test analysis of the proposed model. For evaluating the model’s performance various metrics including accuracy, f1-score, precision, and recall were evaluated. The results show that the Vision Transformer-based model outperformed all the explored models with an overall accuracy of 98% and 97% on the RLVS dataset and the Hockey datasets respectively, which are also significantly higher than the recent solutions proposed on these datasets.