End-to-End Dense Video Captioning Model Based on Multimodal Feature Fusion
Shixin Peng, Ting Xiong, Jingying Chen · 2024
The goal of dense video captioning is to generate multiple text descriptions related to different timestamps from a video. Traditional methods adopt a two-stage "localize-then-describe" approach but overlook the correlation between localization and captioning, leading to inaccurate predictions of event boundaries. Additionally, existing models often rely solely on visual information, neglecting audio and textual subtitle cues, which results in less accurate text descriptions. To address these issues, this paper proposes an end-to-end dense video captioning model. In this model, visual, audio, and textual features of the video are used as inputs to the encoding layer, processed through three structurally identical Deformable Transformer encoders. Three parallel tasks—localization head, captioning head, and event counter—are designed to facilitate end-to-end training. A multimodal feature fusion module is introduced to integrate the three modal features into a temporally coherent feature vector. Finally, the model’s performance is further enhanced through reinforcement learning. Extensive experiments on the ActivityNet dataset demonstrate that the proposed model outperforms PDVC and MDVC in generating higher-quality descriptions.