Attention based CNN-LSTM network for video caption

Xinbo Ai, Qiang Li · International Conference on Mechanisms and Robotics (ICMAR 2022) · 2022

Due to the demand and wide application of video caption in various fields such as video retrieval, content recommendation, risk management, etc., how to extract a comprehensive and highly generalized description of the information has been an active research area for many years. In this paper, we propose a new model that includes the fusion of convolutional neural network and attention mechanism. The special features are extracted from multiple perspectives such as scene, target and behavior in the video, and combined with key frame semantic information to reduce the interference of redundant information and complete the feature representation of the information, and the attention-weighted fusion of the input of the above four features is input to LSTM decoding to finally generate the video content title. A multi-baseline comparison of the two public datasets is performed. Multiple evaluation metrics prove that our model outperforms other models and also show that the information representation of this paper.

Read the paper · More papers on PaperTik