Double Attention Network Based on Sparse Sampling
Zhuben Dong, Yunheng Li, Yiwei Sun, Conghui Hao, Kaiyuan Liu, Tao Sun, Shenglan Liu · 2022 IEEE International Conference on Multimedia and Expo (ICME) · 2022
Locating action segments in long untrimmed videos is a sub-task of video understanding, which more and more scholars pay attention to. Boundary ambiguity and over-segmentation errors are two difficult problems. To handle them, we propose a network called Double Attention Network based on Sparse Sampling (DASS) on the basis of MS-TCN series. First, we design a Seq2Seq Convolution Sampling Network (SCSN) to reduce feature redundancy, which also works on over-fitting. Second, we devise a Global Temporal Attention Module (GTAM) to help predict action boundaries and improve the effect of post-processing from a global perspective. Third, we propose Local Temporal Attention Module (LTAM), which both casts attention to local frames and complements details lost in high dilated layers. We perform experiments on three challenging datasets: 50Salads, GTEA and Breakfast and prove our model is state-of-the-art.