Dual-Scale Alignment-Based Transformer on Linguistic Skeleton Tags for Non-Autoregressive Video Captioning
Xian Zhong, Yi Zhang, Shuqin Chen, Zhixin Sun, Huantao Zheng, Kui Jiang · 2022 IEEE International Conference on Multimedia and Expo (ICME) · 2022
Due to the characteristic of one-time parallel generation of a caption, non-autoregressive video captioning lacks strong dependencies between words. Although using guideline of scene-related visual words can promote caption generation, the semantic relations among visual words are barely explored, limiting the accurate representation. To this end, we propose a Dual-Scale Alignment-based transformer on Linguistic Skeleton Tags (DSA-LST), which alleviates the defect above in the form of visual words group (several words representing a video frame). Different groups represent different semantic dependencies by attention. We utilize linguistic skeleton tags (i.e., several groups) as sentence-level supervision for visual words sequence. For visual words group to accurately express a specific frame, we further design dual scales of visual-language bi-direction alignment to achieve internal relevance of the tags. Extensive experiments conducted on widely used datasets: MSVD and MSR-VTT demonstrate the effectiveness of our method when compared with existing approaches.