Time Efficient Video Captioning Using GRU, Attention Mechanism and LSTM

Gurdeep Saini, Nagamma Patil · 2023

People have a number of video-capturing devices, such as smartphones, cameras, etc. Consequently, they started uploading large amounts of video to the internet, requiring file organization and classification. Many scholars are investing a lot of time in automatically generating sentences for images and videos. It has a variety of uses, such as video indexing, video retrieval, etc. Many approaches are being developed to properly analyze the video and automatically construct the sentence, but understanding the relationship between video information and natural language sentences is still a work in progress. We used the gated recurrent unit, attention mechanism, and long short-term memory unit to develop a recurrent neural network based model for video captioning. We tested our model on Microsoft research video description corpus (MSVD) dataset and showed higher accuracy compared to other state-of-the-art models. The proposed model achieved 0.76 Bleu-1 score and 0.673 Rouge score.

Read the paper · More papers on PaperTik