SPoTem: Handgun Detection in Videos using Spatial, Pose, and Temporal Features
Josue Flores, Liam Liban, Robert Lawrence Ribaya, Idan Andrei Paguio, Jessie James P. Suarez · 2024
Recent advancements in deep learning have led to the creation of automated real-time weapon detection systems, primarily leveraging object detection framework and body pose keypoints extraction to extract features of hand regions as spatial features and body pose features. However, these approaches aren't able to contextualize the spatial features and pose features, limiting their effectiveness. This research introduces a unified architecture called SPoTem that incorporates spatial features, pose features, and temporal features for gun detection. The proposed approach starts by detecting the poses of persons in a video frame using OpenPose to locate the hand regions. Afterwards, the hand region is cropped, and DarkNet-53 is used to extract spatial features from it. Next, human pose keypoints are then put into an image to obtain a binary human pose image which is then fed into a convolutional neural network for human pose feature extraction. As for the temporal component, this study explores two (2) methods of extracting temporal features from the video using a long short-term memory (LSTM). The first one is through using a sequence of pose features and the other is a sequence of spatial features. This study shows that adding temporal features have improved the performance with respect to detecting guns in videos. Furthermore, results have shown that temporal features using sequences of human pose keypoints are more stable and effective at handgun detection in videos.