A Survey of Video Captioning Methods

Rishi Agrawal · 2021 5th International Conference on Information Systems and Computer Networks (ISCON) · 2021

For a human mind, each picture is perceived in a certain way, a different person may see things differently in the same scene. One might say “A dog in grass” while some other might say “A white dog in a grassy area”. All of these captions are relevant and there may be hundreds others too. this seems like a trivial task for human beings, a human being can have a glance and make captions like these with correct use of language with ease. But making a computer do this is a different thing to ask, cognitive ability is required to make captions out of a subjective scene. Without the recent developments in the field of deep neural networks, this would have been an inconceivable task to do, but with these new deep learning neural networks at our disposal, this problem can be solved. Content-based image retrieval, also known as query by image content (QBIC) and content-based visual information retrieval (CBVIR), are different image retrieval methods used to solve the problem of searching for images based on the content of the image. The purpose of this paper is to survey different papers and make a comparison among different methods, algorithms that are useful in video captioning.

Read the paper · More papers on PaperTik