ConDVC: Bridging Visual and Semantic Spaces with Key Semantics for Video Understanding

Jian-Xiong Wang, Jing Xiao · 2024

Dense video captioning task aims to understand unsegmented video content for accurate event localization and captioning. Recent studies have focused on leveraging inter-task relationships to address this task. However, the complexity of cross-modal learning between video content and captions, particularly without prior knowledge guidance, presents significant challenges in jointly handling these two tasks. This paper introduces a Concept-Guided Dense Video Captioning framework (ConDVC) that uses concepts, i.e., key elements such as objects and actions, as a bridge linking visual and semantic spaces. By employing video-to-text retrieval to gather textual features and integrating these with multimodal features for concept detection, we utilize these concepts as semantic guides during the event matching process. This approach not only provides additional prior information for optimizing both subtasks, event localization and caption generation, but also leverages the prior knowledge capabilities of pretrained models like CLIP to enhance overall model performance. Extensive experiments on the YouCook2 and ActivityNet Captions datasets demonstrate the superiority of ConDVC against state-of-the-art methods without extra data for pretraining.

Read the paper · More papers on PaperTik