Movie Caption Generation with Vision Transformer and Transformer-based Language Model
Sorato Nakamura, Hidekazu Yanagimoto, Kiyota Hashimoto · 2023
The paper presents a video caption generation system using Vision Transformer (ViT) and a Transformer-based language model. A video is regarded as images with time information but it is difficult to analyze a video with ordinary image recognition techniques because of the difference between spacial information and time information. So, the system analyzes each image with ViT and integrates all image features into a video feature. ViT is a state-of-the-art object recognition system and consists of stacked Transformers. ViT is not trained with a training dataset but is employed with a pre-trained model. In a caption generation module, a caption is generated with Transformer decoders based on the video feature. The caption generation module is trained with a training dataset. In experiments, we use a large-scale video captioning dataset and train the proposed system. As experiment results, the proposed system achieves 0.27 F-measure in Rouge-L and we confirm that the trained system can generate appropriate captions according to input videos from the viewpoint of human judgment. The results show the proposed system is superior to the system with VGG16.