Hard Contrastive Learning for Video Captioning
Lilei Wu, Jie Liu · 2022
Maximum likelihood estimation has been widely adopted along with the encoder-decoder framework for video captioning. However, it ignores the structure of sentences and restrains the diversity and distinction of generated captions. To address this issue, we propose a hard contrastive learning (HCL) method for video captioning. Specifically, built on the encoder-decoder framework, we introduce mismatchedpairs to learn a reference distribution of video descriptions. The target model on the matched pairs is learned on top the reference model, which improves the distinctiveness of generated captions. In addition, we further boost the distinctiveness of the captions by developing a hard mining technique to select the hardest mismatched pairs within the contrastive learning framework. Finally, the relationships among multiple relevant captions for each video is consider to encourage the diversity of generated captions. The proposed method generates high quality captions which effectively capture the specialties in individual videos. Extensive experiments on two benchmark datasets, i.e., MSVD and MSR-VTT, show that our approach outperforms state-of-the-art methods.