Region-Guided Violence Recognition and Saliency Evaluation in Surveillance Videos

Paris Her, Edwin Engin Yaz, Susan C. Schneider · 2024

Video-based intelligent violence recognition methods should be able to detect the action while also providing explainable results. To that end, this work leverages the location of individuals involved in the action to guide a model’s saliency regions. During the training process, the bounding box regions of the involved suspects are aligned with the model’s salient regions as an additional term in the loss function. Thus, the model inherently learns to recognize these regions. Therefore, during inference no additional computational cost is incurred as this alignment process is discarded. This study presents experimental results on the largest publicly available violence dataset (RWF-2000) to which class labeled bounding boxes have been assigned to individuals. Results demonstrate that incorporating these bounding box regions improves accuracy, sensitivity, and specificity. Model saliency also improves, ultimately leading to better model interpretability. These are the first results on the saliency evaluation for violence recognition in surveillance videos

Read the paper · More papers on PaperTik