Natural Language Description for Videos Using NetVLAD and Attentional LSTM

V. Jeevitha, M. Hemalatha · 2020 International Conference for Emerging Technology (INCET) · 2020

Video captioning infers the process of generating textual description from videos which describes the objects and actions present in it. The multimodal information is available on each frame based on texture and time in the video. In video captioning, the tremendous task is generating the caption automatically related to video content precisely. By using advancement in the field of deep learning technology, a model was developed to generates natural-language descriptions for activities in the video is proposed. In our proposed work, the first stage is extracting the key features for machine understandable about the video content using 2D and 3D CNN. The convolutional neural network(CNN) of 2D and 3D is used to extract both the spatial and temporal features respectively for transferring the videos into key features. The extracted features are preprocessed using NetVLAD. After NetVLAD preprocessing, the features are concatenated and given as input into attention based Long-Short Term Memory(aLSTM). aLSTM generates sentences in a sequential manner by selecting the salient features. The expected output of the model is a sentence to describe the contents of the video. The evaluation is done by using Bilingual Evaluation Understudy (BLEU) metrics.

Read the paper · More papers on PaperTik