Survey of Dense Video Captioning: Techniques, Resources, and Future Perspectives
Zhandong Liu, Ruixia Song · Applied Sciences · 2025
Dense Video Captioning (DVC) represents the cutting edge of advanced multimedia tasks, focusing on generating a series of temporally precise descriptions for events unfolding within a video. In contrast to traditional video captioning, which usually offers a singular summary or caption for an entire video, DVC demands the identification of multiple events within a video, the determination of their exact temporal boundaries, and the production of natural language descriptions for each event. This review paper presents a thorough examination of the latest techniques, datasets, and evaluation protocols in the field of DVC. We categorize and assess existing methodologies, delve into the characteristics, strengths, and limitations of widely utilized datasets, and underscore the challenges and opportunities associated with evaluating DVC models. Furthermore, we pinpoint current research trends, open challenges, and potential avenues for future exploration in this domain. The primary contributions of this review encompass: (1) a comprehensive survey of state-of-the-art DVC techniques, (2) an extensive review of commonly employed datasets, (3) a discussion on evaluation metrics and protocols, and (4) the identification of emerging trends and future directions.