BERTHA: Video Captioning Evaluation Via Transfer-Learned Human Assessment

Luis Lebron, Yvette Graham, Kevin M. McGuinness, Konstantinos I. Kouramas, Noel Edward O'Connor · 2022

Evaluating video captioning systems is a challenging task with multiple challenges to consider.Firstly, the fluency of the caption, multiple actions taking place within a single scene, and estimation of what a human user might consider important in a video.Most metrics aim to measure how similar the system generated captions are to a single or a set of human-generated captions.This paper presents a new method based on a deep learning model to evaluate systems.The model is based on BERT language model, shown to work well across a range of NLP tasks.The aim is for the model to learn to perform an evaluation similar to that of a human.To do so, we use a dataset that contains human evaluation of system-generated captions.The dataset consists of human judgments of the quality of captions produced by the system participating in past TRECVid video to text tasks (Smeaton et al., 2006).These annotations will be made publicly available.The new video captioning evaluation metric, BERTHA, obtains favourable results, outperforming commonly applied metrics in some setups.

Read the paper · More papers on PaperTik