VATEX2020: pLSTM framework for video captioning
Alok Singh, Salam Michael Singh, Loitongbam Sanayai Meetei, Ringki Das, Thoudam Doren Singh, Sivaji Bandyopadhyay · Procedia Computer Science · 2023
Captioning a video involves condensing the video's information into text, which can be useful in video sentiment analysis, video-guided machine translation (VMT), visual question-answering and humanitarian aid. This paper discusses the details of the architecture of the pLSTM framework that is employed for the VATEX-2020 video captioning challenge. In this work, a sequential method is employed wherein to encode visual features a 3D convolutional neural network (C3D) is used. C3D was pretrained using the Sports-1M dataset. In the decoding phase, the input captions and visual features are fused separately in Long Short Term Memory networks (LSTM). The element-wise dot product is performed on the output of both LSTMs to get the final output. On both publicly available and private test data sets, our model achieves BLEU-4 scores of 0.20 and 0.22, respectively.