Video Captioning Using Inflated 3D Convolution Network Encoder with Decoder for Video Content
S. Padmakala, Ramy Riad Al–Fatlawy, Sowmya Madhavan, K. Anuradha, Revathi. R · 2024
Video captioning is a process of automatically generating textual descriptions for video content. This task is crucial in the fields of computer vision and Natural Language Processing (NLP). However, despite recent advancements, there are challenges in capturing temporal dependencies and understanding contextual nuances. To address these issues, a new approach is proposed which combines an Inflated 3D Convolution Network (I3D) encoder and a Long-short term memory (LSTM) decoder for the MSR-VTT (Microsoft Research Video-to-Text) dataset. The I3D encoder extracts spatial and temporal features from video frames and provides visual content representations. These representations are then used as input for an LSTM decoder, which utilizes its long short-term memory capabilities to generate relevant captions. The experimental results on the MSR-VTT dataset prove the effectiveness of the proposed method. The I3D-LSTM approach achieved meteor of $42.56 \%$, bleu- 4 of $46.89 \%$, rouge-L of $\mathbf{6 5. 1 7 \%}$, and Cide of $56.88 \%$ in terms of performance compared to LSTM, ADL (Attention based on a Dual Learning approach), and 2D-CNN (Convolution Neural Network).