Detection of Violent Content in Videos Using Attention-Augmented 3-D Convolutional Networks
Bhavyesh Sajja, Anurag Kumar Singh · IEEE Multimedia · 2025
Many videos over the Internet contain violent and explicit content. As the Internet becomes more accessible, these videos are now available to everyone. It is important to moderate such content as it can have detrimental effects on the mental well-being of people, especially children and teenagers. An automated system can prove useful for the efficient analysis of video content. Inspired by the success of multiheaded self-attention (MHSA) in 2-D CNNs, we propose a novel deep-learning architecture by extending MHSA to 3-D networks. We develop a framework that utilizes the Res3D network and attention-augmented 3-D convolution layers. The proposed model achieves state-of-the-art performance; experiments on the Hockey Fights, the Movie Fights, and the Real Life Violence Situations datasets show accuracies of 98.3 ± 0.6%, 99 ± 1.2%, and 96.85 ± 0.66%. An ablation study was conducted to prove the usefulness of attention augmentation in convolutional neural networks for violence detection.