Video Captioning with Semantic Information from the Knowledge Base
Dan Wang, Dandan Song · 2017
Generating video description is a very challenging task due to the complex spatiotemporal information. Recently, many methods have been proposed by utilizing LSTM to generate sentence for video. Inspired by recent work in machine translation and object detection, we propose a new approach for video captioning which aims to incorporate Knowledge Base information with frame features of the video. We compare and analyze our approach with prior work and show that the large volumes information is available to generate video description. We experiment with our ideas on the S2VT model, and we demonstrate that our method outperforms the state-of-the-art on video captioning benchmarks.