A Deep Learning Framework for Visual to Caption Translation

Anmol Agarwal, Siddharth Garg, Priti Bansal · 2021 3rd International Conference on Advances in Computing, Communication Control and Networking (ICAC3N) · 2021

In reality videography possesses complicated dynamics; and open-domain interpretation methods should be time-sensitive and enable both variable-size input (frame sequence) and output (word sequence). To address the issue, here we suggest a pioneering head-to-head or sequence-to-sequence model for creating video captions or descriptions. We use recurrent neural networks, particularly Long Short-Term Memory (LSTM) networks, to do this, since they have shown avant-garde success in the production of picture captions. The training of the LSTM model is based on video-phrase pairs and it acquires its knowledge by equating a series of visual frames with a succession of words to produce a summary of the snippet's occurrence. The proposed model V2CT has shown signs to grasp both the temporal structure of a series of frames and the series model of the formed phrases, i.e., a dialect model. The model is trained and tested on a common collection of YouTube videos (MSVD dataset).

Read the paper · More papers on PaperTik