Video Captioning Via Two-stage Attention Model and Generative Adversarial Network

Wencai Yan, Lin Qi, Yun Tie, Cong Jin · 2021

Video captioning plays an important role in video retrieval and personalized recommendations with the rapid rise of short videos. In the process of video captioning, the accuracy of caption generation may be affected by the existence of irrelevant visual content. To address this problem, a two-stage attention model is proposed, including object-level attention and frame-level attention. The object-level attention first focuses on the salient objects in each frame and then the frame-level attention adaptively attends to the frame most relevant to the caption. In addition, because of the exposure bias problem in recurrent neural network (RNN)-based sequence generation, generative adversarial network (GAN) and reinforcement learning (RL) algorithm are combined to make the generated caption more accurate and natural. For the discriminator in the GAN framework, we take both the generated caption and the video features as input, then a new reward function is proposed to update the parameters of generator. The experiments on two benchmark datasets, MSVD and MSR-VTT, have shown the effectiveness of our proposed method.

Read the paper · More papers on PaperTik