STD-TR: End-to-End Spatio-Temporal Action Detection with Transformers

Zexian Li, Tian Wang, Aichun Zhu, Kexin Liu, Peng Shi, Hichem Snoussi · 2021 China Automation Congress (CAC) · 2021

Spatio-temporal action detection methods locate human actions in both spatial and temporal dimension, which usually follow a two-stage structure. In this paper, We propose STD-TR, a novel spatio-temporal action detection framework with an end-to-end transformer structure. STD-TR employs two branches to extract feature from video clip and key frame concurrently, then sends the aggregated feature to the transformer encoder-decoder. View spatio-temporal action detection as a set matching and prediction problem, STD-TR employs learned object queries to model the relation of feature context, and directly outputs all predictions at one inference time. Our method remove all hand-designed and can be optimized by a joint loss. Besides, a Hungarian algorithm and a upgraded linking strategy are used for bipartite set matching and action tube generation respectively. Convincing experiment result on challenging dataset demonstrates the superiority of our method.

Read the paper · More papers on PaperTik