A Lightweight Video Surveillance Anomaly Detection Method Based on Image Semantic Representation
Shaohong Li, Zhiguo Shi, Yong Wang · 2024
In this paper, a lightweight method for video surveillance anomaly detection based on image semantic representation is proposed, aimed at enabling widespread deployment on low-cost video surveillance devices. The approach efficiently embeds representations of each video frame and detects anomalies by comparing the similarity between these embeddings and those of normal and abnormal samples stored in a template library. A key innovation of this method is an enhanced lightweight model training technique for image semantic representation, which incorporates specific semantic information into the model's weights, significantly improving anomaly detection accuracy. This approach achieved an impressive 98.7% accuracy on a custom dataset with an inference speed of 9.4fps on the IoT terminal chip K210. Furthermore, an innovative image-to-image generation pipeline for expanding training datasets is introduced. This pipeline generates a diverse set of images with potential anomalies through multimodal understanding and local editing. Additionally, the method features a hardware-friendly and efficient CNN network structure, specifically designed for resource-constrained devices. This structure optimizes parameter efficiency and computational speed, outperforming traditional techniques in both accuracy and efficiency. The proposed method demonstrates superior performance, making it highly suitable for practical deployment in real-world video surveillance applications.