Video Captioning using Pre-Trained CNN and LSTM

A. Preethi, Pandluri Dhanalakshmi · 2023

Digital video is more prevalent nowadays because of more usage of video data among users. The short and catchy videos among social media attract the attention of people. On the same time, the lengthy videos are found to be left without being fully watched. So, video captioning overcomes this issue by automatically generating captions for a video. The process of generating meaningful natural language sentences for the corresponding scenes in the video is called video captioning. Video captioning involves two steps, namely, feature extraction and caption generation. Here, the pre-trained CNN such as InceptionV3 and VGG16 were used for extracting the features from the video. The caption generation is done through LSTM with the help of extracted features. The relevant captions are achieved using LSTM with the help of word embeddings.

Read the paper · More papers on PaperTik