Syntax-Controllable Video Captioning with Tree-Structural Syntax Augmentation
Jiahui Sun, Peipei Song, Jing Zhang, Dan Guo · 2024
Traditional video captioning tasks often suffer from uncontrollable syntax, limiting their adaptability to diverse practical requirements. This paper explores a challenging task known as syntax-controllable video captioning (SCVC). Specifically, this task aims to generate captions for videos that convey the visual content accurately and adhere to the syntactic structure provided in an example sentence. To tackle the SCVC task, we propose a novel Tree-Structural Syntax Augmentation Network (TSAN). It consists of a tree-structural syntax encoder to capture syntactic parent-child relationships of example sentences and a syntax-customized caption generator to integrate syntax and semantic hints effectively. Moreover, we introduce a syntax-semantic joint optimization strategy that considers both syntax and semantic constraints for joint training. Considering the fact that there are no available syntax-controllable video caption datasets, we use the traditional video captioning dataset to train the model. In the inference process, the model generates syntax-controllable captions in a zero-shot manner, applying exemplar sentences randomly selected from an open corpus and unseen in the training data. The experiments on MSRVTT and VATEX datasets validate the superior performance of TSAN in both semantic and syntactic evaluations. Extensive ablation studies and visualizations further show the effectiveness of the proposed method.