Research on Video Captioning Method Based on Semantic Key Frame

Yan Li, Xu Cui, Xiaofeng Jin · 2022 2nd Asia-Pacific Conference on Communications Technology and Computer Science (ACCTCS) · 2022

At present, in most video captioning models, when extracting video key frames, the extracted video key frames not only contain a lot of redundancy resulting video frame features with small diversity, but also ignore the video semantic information. Due to the neglect of such uniqueness, a semantic based video key frame method is proposed in this paper. Firstly, the initial video frames are extracted by combining the depth convolution network(CNN) and feature windowing. Secondly, this paper proposes to use the image captioning network to annotate the video frames and find the video key frames according to the semantic information. This method extracts the video key frames with more diversity and low redundancy. Then, we extracts the spatial, dynamic and regional features of the video to capture the potential context information in the video. Finally, we propose a pre-interactive LSTM model to effectively integrate text features, video features and context features, so as to solve the problem of context information loss and improve the annotation performance. Through the comparative experiment on MSR-VTT datasets, the results show that the proposed method improves by 1.2, 0.2, 0.1, 0.3 and 1.5 compared with the baseline model in terms of BLEU4, METEOR, ROUGE-L and CIDEr respectively, the MSVD datasets is improved by 2.2, 0.5, 0.4 and 4.4, which effectively improves the performance of video captioning.

Read the paper · More papers on PaperTik