Video Captioning using a Hybrid Transformer and RNN-based Encoder-Decoder
Alexandru-Cosmin Mihai, Mihai Masala, Dan-Teodor Poncu, Traian Eugen Rebedea · 2022
Video captioning refers to generating a textual representation of the content presented visually in a video, a task that is necessary both for visually impaired users and for a better navigation through large video repositories.In this paper we describe a model that employs the heavily pretrained CLIP Vision Transformer (ViT) [9,5] and the GPT2 language model [10] and discuss several approaches of combining the two models for video captioning.Our method relies on a novel hybrid model that combines the pretrained CLIP ViT with LSTM networks, outperforming CLIP + GPT2 models.We validate our method on the popular MSVD [4] and MSR-VTT [18] datasets, showing the potential of the proposed model for video captioning.