Critic-based Attention Network for Event-based Video Captioning

Elaheh Barati, Xuewen Chen · 2019

In this paper, we investigate utilizing an actor-critic architecture for an event-based video captioning. In this captioning task, the video can contain multiple overlapping events. We consider the words as the actions that our model takes sequentially. The architecture of our model consists of an actor network to predict the captions given temporal segments in a video and a critic network to measure the quality of the generated captions. Our model, first, localizes events in the video, then by using the localized events, it locates temporal segments in the video. We adopt a global network to generate a caption for each temporal segment. We propose an attention mechanism to account for the importance of each localized event in captioning a temporal segment in the video. We provide a set of experiments on utilizing our method in the task of event-based video captioning on ActivityNet Captions and TACoS-MultiLevel datasets. Experimental results show that our method outperforms state-of-the-art video captioning methods.

Read the paper · More papers on PaperTik