The Video Captioning Method Based On The Spatial- Temporal Information and Attention Mechanism
Ou Ye, Tao Liu, Yan Fu, Jun Ping Deng, Jian Feng, Yun Zhang · 2021
In order to utilize the complementarity of different level features in different regions of videos effectively, and improve the accuracy of text description in videos, we propose a video captioning method based on the spatiotemporal information and attention mechanism. First, the Faster-RCNN and VGG-16 networks are used to extract the high-level features of the interesting regions and significant targets in the video frames, respectively. Second, combing the low-level dense traj ectory features of video frames to represent the spatial appearance feature of a video. Finally, we adopt the long short-term memory network to extract the temporal association features between frames and texts. In the hidden layer of the long short-term memory, an adaptive attention mechanism is introduced to enhance the representation of the salient features for video frames, and then the text description is generated automatically. The experimental results show that our proposed method can effectively improve the accuracy of video captioning, and its performance is better than several the state-of-the-art approaches.