Attention-based LSTM with Semantic Consistency for Videos Captioning
Zhao Zhuang Guo, Lianli Gao, Jingkuan Song, Xing Xu, Jie Shao, Heng Tao Shen · 2016
Recent progress in using Long Short-Term Memory (LSTM) for image description has motivated the exploration of their applications for automatically describing video content with natural language sentences. By taking a video as a sequence of features, LSTM model is trained on video-sentence pairs to learn association of a video to a sentence. However, most existing methods compress an entire video shot or frame into a static representation, without considering attention which allows for salient features. Furthermore, most existing approaches model the translating error, but ignore the correlations between sentence semantics and visual content.