A Framework For Captioning The Human Interactions
U Afreen Farzana, S. Abirami, Mrs. J. Srivani · 2019
Caption generation is an emerging Artificial Intelligent challenge where a content description has resulted in a given input. Captioning involves the Computer Vision methodologies for the identification of content from input images and language modeling techniques for processing the text. The objective of Video Captioning is to generate a natural language sentence relevant to the content of the input video clips. In this paper, a deep learning-based encoder-decoder model has been used to result in effective video captions for human actions. The Caption Generative model takes video as input and generates a caption for the interactive actions performed by a human. This model comprises of two stages. The first stage (Encoder) performs extraction of the features using the Inception V3 model in Convolution Neural Network (CNN), and the second stage (Decoder) uses Long Short Term Memory (LSTM) a sequence modeling neural network is used for generating the captions. SBU Interaction dataset is used to evaluate the framework dealt in with this paper. Metrics such as accuracy, recall, precision, and F-score are measured to demonstrate the performance of the model. Bilingual Evaluation Understudy (BLEU) Score is also calculated for evaluating the generated captions.