A Method for Dense Video Captioning based on Action Proposals
Ruoshan Sun · 2025
Dense video captioning aims to localize multiple events in one or more unedited videos and generate detailed video descriptions of these events. Current end-to-end dense video captioning models mainly use a video description framework based on object detection algorithms, which focuses on the classification labels of objects and suffers from insufficient learning of action labels. In this paper, we propose a Dense Video Captioning with Action proposal algorithm, which generates action proposal queries based on temporal action detection in addition to event queries for the current video. The experimental results show that compared with the reference algorithm, the proposed algorithm improves about 3%, 1.7%, and 2.8% in the BLEU@4, METEOR, and CIDEr metrics, respectively, for the dense video captioning task on the Activity-Net public dataset.