Joint Spatio-Temporal Action Localization in Untrimmed Videos with Per-Frame Segmentation

Xuhuan Duan, Le Wang, Changbo Zhai, Nanning Zheng, Qilin Zhang, Zhenxing Niu, Gang Hua · 2018

Inspired by the recent spatio-temporal action localization efforts with tubelets (sequences of bounding boxes), we present a new spatio-temporal action detector Segment-tube, which consists of sequences of per-frame segmentation masks. The proposed Segment-tube detector can temporally pinpoint the starting/ending frame of each action class in the presence of preceding/subsequent interference actions in untrimmed videos. Simultaneously, the Segment-tube detector produces per-frame segmentation masks instead of bounding boxes, offering superior spatial accuracy to tubelets. This is achieved by alternating iterative optimization between temporal action localization and spatial action segmentation. Experimental results on multiple datasets validate the efficacy of the proposed detector.

Read the paper · More papers on PaperTik