SF-CDDM: Skeleton Feature-Driven Conditional Denoising Diffusion Model for Video Anomaly Detection
Siyu Feng, Hong Xia, Hui Jia, Yanping Chen · 2025
Video anomaly detection plays a critical role in video surveillance systems, particularly in applications related to public safety and industrial production. However, existing detection models still face challenges in capturing complex spatiotemporal dynamic features and accurately identifying anomalies. This paper proposes a diffusion model based on skeletal features, incorporating spatiotemporally separable graph convolution (STS-GCN) and attention mechanisms to improve anomaly detection accuracy. First, the DDIM framework is used to add noise to the skeletal features, and a spatiotemporally separable graph convolution autoencoder is employed to extract spatiotemporal dependencies in the skeletal features. After integrating temporal information, the features are passed into an STS-GCN-based U-Net for denoising. The model introduces a self-attention mechanism within the autoencoder to optimize the representation of spatiotemporal features, and employs cross-attention mechanisms to fuse multisource information, enhancing both denoising and anomaly detection accuracy. Experimental results demonstrate that the proposed model outperforms existing methods on several benchmark datasets (such as HR-UBnormal, HR-STC, and HRAvenue), validating its effectiveness and innovation.