A Deep Learning System for Deep Surveillance

Aman Anand, Rajendra Kumar, Nikita Verma, Akash Bhasney, Namita Sharma · 2024

Deep surveillance is an essential task in computer vision that involves monitoring and analyzing video data to detect and track objects of interest, identify unwanted events, and ensure public safety. Analysis of real-time data has gained popularity in the last 10 years due to its remarkable results analyzing videos and interpreting pictures in many ways. By combining a variety of low-level picture characteristics with comparatively high-level information from object detectors and scene classifiers, they may easily stagnate their performance. This chapter presents a deep learning model with different implementations of convolutional neural networks (CNNs) for deep surveillance applications. The proposed model leverages the power of SoftMax Regression, Support Vector Machine, Convolutional Neural Network, MatConvNet, and Spatially-sparse CNN to achieve robust object detection, tracking, and anomaly detection from real-time video streams. The model incorporates both spatial and temporal information for comprehensive analysis and integrates various architectural innovations for improved performance. The mathematical model uses coordinate hashing for mapping the coordinates in different hierarchical layers of the convolution process for object detection. The Hash function uses input coordinate, stride, dilation, neighbor counting which cannot be zero, and the size of the dimension of the feature map. The coordinate map is generated as a dictionary stored in memory. Transpose convolution and pruning of coordinate map are used for cutting off the extra generated coordinates to make the layer learnable and model trainable with high accuracy. While implementing the mathematical model for object detection, for pre-processing, the video is converted into frames and then analyzed using a pre-trained deep-learning model to analyze for deep surveillance. The extracted frames (or feature maps) are passed through the filter (kernel) and MaxPooling layer for reduction of size. The small size of the feature map reduces the computation cost. As this study presents model variants for deep-leaming-based object detection for analysis of the security of human beings, the Spatially-sparse CNN implementation observed the best performance in terms of training accuracy and object detection results.

Read the paper · More papers on PaperTik