Dense Video Captioning through Convolutional-Transformer Integration

Jia‐Xi Cao, Chunjie Zhang, Xiaolong Zheng · 2024

Dense video captioning, a task that has recently garnered attention in the field of computer vision, focuses on generating detailed captions for multiple events along with their temporal locations within untrimmed videos. Existing transformer-based methods have achieved notable success. However, some methods utilizing deformable transformer face challenges in processing long videos, as the encoder may struggle to capture the long-term dependencies of events in the video. This issue is partly due to not integrating Convolutional Neural Networks (CNN) with the deformable transformer architecture. In this work, we introduce Dense Video Captioning through Convolutional-Transformer Integration (CTDVC), which links Layer Normalization, the Scalable-Granularity Perception (SGP), and Group Normalization into a network, replacing the encoder part of the deformable transformer. This model better accounts for the sequential relationship between events, especially those events in the video that span over a longer duration, allowing it to generate video captions that are not only more coherent but also more precise in their narrative. We evaluated our method on ActivityNet Captions and YouCook2 datasets. Experiments demonstrate that our proposed method CTDVC achieves an improvement in performance.

Read the paper · More papers on PaperTik