Multi-Perspective Video Captioning

Yi Bin, Xindi Shang, Bo Peng, Yujuan Ding, Tat‐Seng Chua · 2021

This work targets at the problems of comprehensive video captioning and the generation of multiple descriptions from different perspectives, termed asMulti-Perspective Video Captioning. We build and release a dataset named VidOR-MPVC, the first dataset for multi-perspective video captioning, where each video is annotated with multiple descriptions from different perspectives. We also propose a novel model, dubbedperspective-aware captioner (PAC), which is capable of mining the various perspectives in a video and generating a description from each perspective. More specifically, a perspective generator is designed to perceive video content with perspective preferences, and followed by a language generator equipped with perspective-aware attention mechanism. As our new task expects to produce multiple descriptions for a video, existing evaluation metrics are fail to handle this situation. To address this problem, we devise the maximum matching scores based on existing metrics for an overall evaluation which aims to cover the aspects of semantic similarity, completeness and compactness. The experimental results demonstrate that our model is able to describe videos with multiple descriptions from different perspectives.

Read the paper · More papers on PaperTik