Advanced Detection of Violence From Video: Performance Evaluation of Transformer and State of the Art of Convolution of Neural Network Transformer
Abdulrahman Alshalawi, Wadood Abdul, Ghulam Muhammad · IEEE Access · 2025
Safety is paramount in every aspect of human concern. In situations where violence occurs beyond the control of observers, it is essential for trained professionals to intervene. These scenarios can be detected by advanced surveillance systems equipped with cameras or sensors. The performance of these systems is highly dependent on the models they incorporate. Although various machine learning and deep learning models have been explored to address these challenges, significant work remains in handling a diverse range of images. This study aims to bridge this gap by training four models—VGG16, MobileNetV2, YOLOv6 and Ensemble Model, and a vision Transformer model with tuned parameters on publicly available video frames. Our analysis established that the Transformer model achieved the highest accuracy, making it the most suitable for applications involving the detection of violent incidents. This research introduces a novel application of AI for semi-automated violence recognition, demonstrating AI's significant potential in improving public safety. It also emphasizes the superior performance of the Transformer model in accurately detecting violent frames. Semi-automated violence recognition refers to a system where AI models assist in identifying violent incidents within video footage, but human intervention is still required for certain tasks, such as verification or decision-making. The AI system automatically detects potential violent events, flagging them for further review by human operators, who then confirm or act upon these alerts. The Transformer model achieved outstanding results: for the Violence Dataset, it attained a precision of 99%, recall of 99%, F1 score of 99%, and accuracy of 99%. To further verify the performance of the Transformer model, we tested it on the Road-Anomaly Dataset, which presents a challenging scenario with imbalanced classes. For this dataset, the Transformer model achieved a precision of 98%, recall of 97%, F1 score of 98%, and accuracy of 98%. These results affirm the model's efficacy, positioning it as the top choice for real-time surveillance and public safety applications.