Indoor Violence Detection using Lightweight Transformer Model
Arushi Kumar, Arpan Shetty, Archit Sagar, A Charushree, Preet Kanwal · 2023
Human activity recognition from surveillance videos is an active research area in Computer Vision and Machine Learning. The detection of indoor violent activities is a major challenge in the surveillance sector due to the lack of sufficient indoor footage, and the presence of obstacles inducing occlusions in a household environment. Several studies on the topic of violence detection have been done, with the best performing models using Convolutional Neural Networks for spatial feature learning, followed by a Recurrent Neural Network variant for temporal feature learning. Recent studies suggest that Vision Transformers have a better run-time and are more robust to partial occlusions than their deep convolutional network counterparts. In this paper, we present a light-weight transformer model, drawing upon the recent success of video vision transformers in action recognition. The model extracts spatial features from the input frames, and adds temporal relation to the selected frames with the help of tubelet embedding, which are then encoded using various transformer layers. Generally, although transformer models are shown to need extensive training on large datasets, it has been proven that using efficient preprocessing, we can train the model on comparatively small sized datasets. The light-weight model has achieved an accuracy of 88.24% on our self-curated indoor violence dataset and a 98% accuracy on a subset containing videos in which the subjects involved in the activity are occluded from view.