Human Keypoint Change Detection for Video Violence Detection Based on Cascade Transformer

Li Zhou, Wenjie Li, Yu Chen, HanLing Liu, MentTing Yang, ZiLong Liu · 2023

Human action recognition remains a challenging task in computer vision. We focus on violence detection, a subset of action recognition. In this paper, we present a video violence detection model, which introduced cascaded Transformer for human pose estimation and 3D-CNN for capturing motion present in spatial-temporal dimension. The cascaded Transformer consist of person detecting Transformer and human keypoints detection Transformer. We trained 3D-CNN network to learn spatialtemporal features from human keypoints sequence, and output whether violence is present in video. We analyzed the influence of frame amount and human keypoints connection way on motion feature extraction. The proposed model achieves an accuracy rate of 83.2% on the RWF-2000 validation set and 90.9% on our own campus violent video dataset.

Read the paper · More papers on PaperTik