Enhanced Video Caption Model Based on Text Attention Mechanism

Yan Ma, Fang Meng · 2022 5th International Conference on Data Science and Information Technology (DSIT) · 2022

In the current video caption model, there are two main problems: only the word of the most recent previous time step is used to represent the contextual information, resulting in insufficient expression of text information affecting the quality of the generated captions; and the use of the word-level cross-entropy loss optimization model makes the training and test inconsistent, which may easily lead to the problems of exposure bias and insufficient model optimization. Considering the above two problems, we first propose a kind of attention mechanism based on historical textual information. Soft selection of the generated historical classification distribution is guided by using the current text encoding to enhance contextual information on the current moment. Then the maximum reward via the multi-sampling method is used to improve the strategic gradient algorithm, which exploits a sentence-level loss to directly optimize the evaluation metric. The Experiments on the Microsoft Research Video Description Corpus (MSVD) dataset shows that our proposed method is comparable with state-of-the-art algorithms.

Read the paper · More papers on PaperTik