Context Visual Information-based Deliberation Network for Video Captioning
Min Lu, Xueyong Li, Caihua Liu · 2021
Video captioning automatically and accurately generates a textual description for a video. The typical methods following the encoder-decoder architecture directly utilize hidden states to predict words. Nevertheless, these methods do not amend the inaccurate hidden states before feeding those states into word prediction. This leads to a cascade of errors in generating word by word. In this paper, the context visual information-based deliberation network is proposed, abbreviated as CVI-DeINet. Its key idea is to introduce a deliberator into the encoder-decoder framework. The encoder-decoder first generates a raw hidden state sequence. Unlike the existing methods, the raw hidden state is no longer directly used for word prediction but is fed into the deliberator to generate the refined hidden state. The words are then predicted according to the refined hidden states and the contextual visual features. The results on two datasets show that the proposed method significantly outperforms the state-of-the-art methods.