MVC: Multi-stage video caption generation model based on multi-modality

Wendong Zhang, Pu Sun, Lan Peng, Zhifang Liao · 2025

In image caption generation, embedding image captions as a feature into the model input has been proven to be an effective method. However, in the field of video captioning, the input features of most models are still single-modal videos without text features; in addition, the captions generated by current video captioning models are relatively simple and not specific and detailed enough.To solve these problems, this paper proposes a new MultiModal Coarse-to-Fine caption generation framework, which consists of three parts: a retrieval module, a coarse-grained caption generation module and a caption polishing module. We conducted extensive experiments on the MSR-VTT and MSVD datasets, and the experimental results showed that MVC performs comparable to or even better than the state-of-the-art methods in terms of CIDEr and METEOR metrics.

Read the paper · More papers on PaperTik