Video Captioning model using a curriculum learning approach

Soumya Varma, James Dinesh Peter · 2022

Video captioning techniques have wide applications in a plethora of domains including surveillance, assistance of visually impaired, object tracking and much more. It is the process of generating a meaningful description about a video in a natural language without losing the semantics of the content in the video. Various works in this area have been thriving since last decade and has reached up to the usage of Transformer networks. Little significance has been made on the quality of the data used for training the model. The captions of the video may be of variable length and if not wisely trained, it can create a biased model which is noisy and irrelevant. In order to overcome this problem, we make use of an improved curriculum learning methodology in video captioning system which is very promising as per the experiments conducted so far. The Curriculum learning strategy which mimics human learning process, can rank the training data according to the difficulty measurement and hence the training model get appropriate data suitable for the training phase. The datasets and metrics used are being evaluated with some standard model and the CL approach turns to be promising.

Read the paper · More papers on PaperTik