Video Captioning using Spatio-temporal Graphs: An Encoder-Decoder Approach
Vibhor Sharma, Anurag Kumar Singh, Sabrina Gaito · 2024
To understand the content in the video, describing it in the form of natural language is also one of the essential requirements. A video consists of frames showing different objects and actions. These objects can change in each frame, and there is a time-dependent sequence between objects. Objects in the sequence of frames have a temporal relationship. Actions in a frame are the interactions between objects, which can be easily understood using spatial relationships. When spatial and temporal aspects combine, it is the Spatio-temporal relationships between objects. Analyzing this relationship is done using a Spatio-temporal graph, which integrates both visual and motion features of detected objects. An integrated Spatio-temporal graph describes the video in a natural sentence. The generated sentence is evaluated using Metrics like BLEU@4, METEOR, and CIDEr.