A Tale of Two Modalities for Video Captioning
Pankaj Joshi, Chitwan Saharia, Vishwajeet Singh, Digvijaysingh Gautam, Ganesh Ramakrishnan, Preethi Jyothi · 2019
Recent advances in machine learning have led to significant accuracy improvements for the task of generating textual captions from videos based on audio and visual signals. In this work, we focus on the influence of modality (audio and visual input) on semantic coherence and well-formedness of the generated captions. We explore both architectural and algorithmic choices that potentially influence the utilization of these modalities. (Algorithmic choices include pretraining while architectural choices include modality-specific weighing schemes.) We study the influence of our choices on a popular video captioning dataset, MSRVTT, by providing quantitative and ex-tensive qualitative evaluations that measure the influence of audio-visual modalities, cohesiveness and ranked relevance of keywords. We are able to assert qualitative improvements on metrics characterizing the quality of the captions, along with obtaining performance that is comparable to state-of-the-art on standard quantitative metrics such as BLEU-4, METEOR, etc.