Research on Multi-model Mask Video Caption based on Transformer

Yuqing Peng, 昊 姚, Xinhao Ji, Xuan Gao · Research Square · 2023

Abstract Previous work on video caption has mainly focused on visual information in the video, neglecting the audio information. This differs from how humans perceive the world by simultaneously processing and integrating multiple modalities. This paper proposes a novel multimodal masked video captioning model (MMVC) that uses a new Transformer model to integrate both the audio modality, which contains rich semantic association information, and the masked video frame image modality. By studying different fusion stages and strategies, a new fusion strategy is proposed that improves model performance while reducing the computational cost. Experiments on the MSVD and MSR-VTT datasets show that the audio modality contains complementary information that effectively improves video caption performance and metrics.

Read the paper · More papers on PaperTik