Exploring adaptive attention in memory transformer applied to coherent video paragraph captioning
Leonardo Vilela Cardoso, Silvio Jamil F. Guimarães, Zenilton K. G. Patrocínio · 2022
A coherent description is the ultimate goal regarding video captioning via a couple of sentences because it might also affect the consistency and intelligibility of the generated results. In this context, a paragraph describing a video is affected by the activities used to both produce its specific narrative and provide some clues that can also assist in decreasing textual repetition. This work proposes a model, named Adaptive Transformer, that uses attention mechanisms to enhance a memory -augmented transformer. This new approach increases the coherence among the generated sentences, assessing data importance (about the video segments) contained in the self-attention results and uses that to improve readability. The test results show the potential of this new approach as it provides higher coherence among the various video segments, decreasing the repetition in the generated sentences and improving the description diversity in the ActivityNet Captions dataset.