Multi-loss Spatial-Temporal Attention-Convolution Network for Action Tube Detection

Jinlei Zhu, Houjin Chen, Pan Pan, Jia Sun, Kun Jing, Chuanfeng Zhang · 2021

The paper proposes a novel network model of the video action tube detection method based on a 3-dimension convolutional neural network and spatial-temporal attention mechanism. The network introduces a special module for extracting the location features of the video sequences, and an internal loss function is designed for the location network module. The location features are then squeezed to modify the scale to be the same as the classification network feature. They are then concatenated and input into a fusion network that receives feedback from the global loss function. The overall position deviation of the target in the video sequence is taken as the internal loss. This directly affects the squeezed location feature of the sequence, and improves the accuracy of action prediction under the attention mechanism. Experimental results indicate that our network model outperforms the existing methods. We further visualize the activation map which reveals the intrinsic reason for performance improvements.

Read the paper · More papers on PaperTik