Object Centered Video Captioning using Spatio-temporal Graphs

Vibhor Sharma, Anurag Kumar Singh, Sabrina Gaito · 2024

Video data is increasing very rapidly day by day. A video description helps to understand the content of the video. Video Captioning is a term used to describe a video in a human understandable language. A video is a sequence of frames. Each frame consists of particular objects. Since, there may be one or more than one object in a frame, objects within a frame can be isolated or may have some interaction with other objects. The interaction between objects with in the frame is the relation between objects. Particular objects may appear, disappear, or change in the sequence of frames. The two objects may be related to each other, or one may perform a specific action with other objects. In other Scenario, one object may be affected by other object. In both of these conditions, objects are the center point of the scene. So, an object-centered approach is a most important choice to generate natural language sentences corresponding to the video. Objects and their relations both are considered for mapping with the graph. A graph is the most essential choice to trace objects across the frame. A graph with space and time maps all objects across the frames in a video. Objects and spatiotemporal graphs are used to generate natural language sentences. Experiments are done on the Microsoft Research Video to Text (MSR-VTT) data set. Object-Centered video captioning approach using a Spatio-temporal approach evaluates the score in terms of evaluation metrics like BLEU@4, METEOR, and CIDEr.

Read the paper · More papers on PaperTik