Generating Natural Video Descriptions using Semantic Gate
Hyungmin Lee, Il-Koo Kim · 2019
Video captioning task aims to generate a textual description of the situation in a video. It is challenging because of the nature of modality-difference between video and language. We present a novel method to bridge the gap between them by utilizing the semantic gate in two ways. First, we develop an activation mechanism to make a video description that captures the concept of the video. Next, we design a network that evaluates the similarity between visual and sentence feature. Semantic gate is used to transform sentence into a semantic embedding. We also conduct experiments to show that image and action classification task performance is transferred to video captioning task. Experimental results show that our proposed method has gained promising improvements compared to the baseline model. Consequently, our model demonstrated the effectiveness by achieving new best record on MSRVTT and MSVD dataset.