Semantic Knowledge Distillation for Video Captioning

Huilan Luo, Siqi Wan, Xia Cai · 2024

Video captioning is a complex task. Most recent works have designed novel models to improve the effect of video captioning generation. However, these methods do not balance the computational cost and model performance well, leaving a lot of room for improvement. This time, we consider the video captioning task by combining the knowledge distillation method and design a dual-branch network framework SFTCap that uses a teacher network to optimize the student network. Compared with other video captioning methods, our SFTCap has the following three advantages: 1) We design a new encoder-decoder based network to understand the video content and decode different levels of semantic information representation; 2) We propose semantic knowledge distillation: using a complex teacher network to improve the student network's ability to capture key semantic information in the video. Specifically: the teacher network passes the object and action semantic representation of the video content it captures as privileged information to the student network, combined with our designed training strategy, to achieve the effect of optimizing the student network; 3) Compared with other video captioning tasks that use additional object detectors and 3D features, we only use the visual features of the video for modeling and training. Experimental results on the challenging MSR-VTT and MSVD datasets show that our SFTCap outperform recent advanced methods.

Read the paper · More papers on PaperTik