Enhancing Ocean Scene Video Captioning with Multimodal Pre-Training and Video-Swin-Transformer
Xinyu Chen, Meng Ya Zhao, Fan Shi, Meng'en Zhang, Yu Hua He, Shengyong Chen · 2023
With the success of multimodal pre-training models in the video-language field and various downstream tasks, previous multimodal models used 3DCNN networks as video feature extractors, which have limitations in interacting and fusing with text features. This paper proposes a multimodal pre-training model that utilizes a Video-Swin-Transformer-based network to encode both video and text data, to achieve better performance in video understanding. The model consists of four modules: video encoder, text encoder, interact encoder, and caption decoder to accomplish the task of ocean scene video captioning. A dataset of ocean scene videos, including various content types such as sea surfaces and shores, is also constructed. The training process is divided into two stages: pre-training and fine-tuning. Pre-training is performed on the Howto100m dataset to allow the model to learn video captions in natural scenes and complete video-language matching tasks. The fine-tuning stage is then performed on the ocean1000 dataset to better understand the events and content in ocean scene videos and generate captions that conform to ocean scene video descriptions. The model achieves satisfying results on both the public dataset YouCook2 and the proprietary dataset Ocean1000, demonstrating its ability in video-text information fusion and interaction.