A Novel Method for Sequence Generation for Video Captioning by Estimating the Objects Motion in Temporal Domain

Prashant Kaushik, Vikas Saxena, Amarjeet Prajapati · 2024

For disruption in video based ai tasks, we need better understanding of features and their sequences. Sequence generation for video captioning has been a tedious task for better and relevant feature mapping for tasks like mapping to objects or text generation. It has gone from sequence to selective sequence, merged sequences, then reweighted sequence based on sematic positions of objects in video. Mostly the sequence is objects position detected in frames & aligned as sequence and merged sequence in based on some event. This paper is estimating the relative motion of object as path. These paths are merged to generate a path of group of objects, and converted to sequence in temporal domain. This way generating the sequence which are better and semantically similar to the ground truth of the text of video for object-oriented video captioning datasets. To prove the workability of methods, some synthetic videos with few objects and background blurring with GMM is used. Lastly, we have calculated and shown the similarity of the generated sequence to the ground truth of the video, and compared it with the sequence generated by the earlier methods of the sequence generation with various features. The comparison of the sequence is done on 4 kinds of videos like synthetic, real videos and the result paths are also shown as plots of the comparison.

Read the paper · More papers on PaperTik