Incremental Information-Aware: Mine Abundant and Accurate Information for Video Captioning
Ningkai Zhong, Bin Fang, Mengdi Li, Langping Wang · 2025
This paper proposes the Incremental Information-Aware (IIA) framework to enhance video captioning by generating semantically abundant and accurate descriptions. The Semantic Incremental Information Aware model (Semantic-IIA) aims to capture detailed semantic information, while the Structural Incremental Information Aware model (Structural-IIA) focuses on identifying key structural content. These models are designed in parallel to complement each other, ensuring both semantic abundance and structural accuracy in generated captions. Experiments on the MSVD and MSR-VTT datasets demonstrate superior performance, with CIDEr scores of 108.4 and 59.8, and BLEU-4 scores of 61.0 and 47.3, respectively. Our code is available at https://github.com/Zhongnibug/IIA.