Cross-frame Reverse Attention for Video Captioning
Huilan Luo, Xia Cai, Siqi Wan · 2024
Video captioning aims to generate natural language descriptions for the content presented in the video. Due to the variety of content presented in the video, the variety of people and objects appearing in the video, it is easy to detect the objects appearing in the video, but it is difficult to fully explore the spatial and temporal relationships between objects and determine the main objects of the events in the video. To solve the above problems, in this paper, a new attention mechanism is adopted to generate video captions. CRACaps is based on the encoder-decoder architecture. It pays attention to the visual representation of the main object in each video frame through attention, and reasons the visual relationship between the object of the first frame or key frame and the insignificant target objects in other frames through the reverse attention mechanism. While building the relationship between objects, it learns the main object information and the action change information of the main target object. We have conducted a large number of experiments on the MSVD and MSR-VTT datasets, and the experiments show the superiority of our method over the state-of-the-art methods.