Confiner Based Video Captioning Model

J Vaishnavi, V. Narmatha · IOP Conference Series Materials Science and Engineering · 2022

Video captioning is the interesting task of encoding features and decoding the encoded features into natural language. Video captioning task is the perfect blend of computer vision and Natural language processing techniques. Video captioning process is a tentative mission since it must consider temporal features along with spatial features in order to generate appropriate captions. This task has the default framework as encoder-decoder. Ensuring the optimality of the generated caption for the particular video is the most important and essential thing which is not considered by most of the existing works. In order to ensure the optimality of captions for videos, our proposed work introduces confiner to the default framework. Confiner is planned to diminish the semantic gap among the generated caption and videos. Confiner is designed in our model with both LSTM and GRU separately as two different models. The actual work of the confiner is to inspect the visuals of the generated caption with the actual video. The proposed model has experimented with three different benchmark datasets as MSVD, MSR-VTT, and M-VAD. The performances of the models are measured by BLEU, METEOR, CIDEr evaluation metrics.

Read the paper · More papers on PaperTik