Video Captioning Method Based on Semantic Topic Association

Yan Fu, Ying Yang, Ou Ye · Electronics · 2025

To address the issue in encoder–decoder-based video captioning method where semantic topics are overlooked in guiding the video captioning, leading to a semantic mismatch between video content and generated captions, a video captioning method based on semantic topic association is proposed. The core of this method lies in the designed semantic topic module, which is capable of capturing more detailed topic information on objects and actions in videos, and establishing a bridge between video representations and semantic topics, thereby selectively guiding the generation of video descriptions. During the decoding phase, an optimized NGPT kernel sampling decoding strategy is employed, constructing an NGPT model based on Top-P kernel sampling to jointly decode semantic information and generate captions. This model fully leverages the topic information in the video, narrowing the “semantic gap”, mitigating redundancy issues during decoding, and making the generated text descriptions that are more aligned with the topic content of the video. Ultimately, experiments were carried out using two publicly accessible datasets: MSVD and MSR-VTT. The findings reveal that our proposed approach surpasses several leading methods, achieving enhancements of 5.3% in BLEU4, 3.1% in METEOR, 2.4% in ROUGE_L, and 13.1% in CIDEr on the MSVD dataset, as well as improvements of 0.9%, 0.4%, 1.5%, and 3.3%, respectively, on the MSR-VTT dataset. Therefore, the video captioning method based on semantic topic association can extract more topic information, thereby improving the effectiveness of video caption generation.

Read the paper · More papers on PaperTik