Single-Stage End-to-End Temporal Activity Detection in Untrimmed Videos
Md Atiqur Rahman, Robert Laganière · 2020
In this paper, we address the problem of human activity detection in temporally untrimmed long video sequences, where the goal is to classify and temporally localize each activity instance in the input video. Inspired by the recent success of the single-stage object detection methods (e.g., SSD [1]), we propose an end-to-end trainable framework that learns task-specific spatio-temporal feature representations of a video sequence using a 3D convolutional neural network followed by a multi-scale detection architecture for direct classification and localization of the activities at varying temporal scales. Our proposed approach sets new state-of-the-art on the highly challenging THUMOS’14 temporal activity detection bench-mark – up from 44.2% mAP to 49.0% mAP (an absolute 4.8% improvement) when the tIoU threshold is set to 0.5 during evaluation.