Bi-Directional Self-Attention with Relative Positional Encoding for Video Summarization

Jingxu Lin, Sheng-hua Zhong · 2020

Video summarization technique is a promising approach to process large-scale video by shortening the video content into a compact version. Most previous methods used recurrent networks such as Long Short-Term Memory (LSTM) even combined with attention mechanism to achieve state-of-the-art results. However, these networks are complex to implement and cannot be easily parallelized in computation. In this paper, we propose a novel deep summarization framework named Bi-Directional Self-Attention with Relative Positional Encoding for Video Summarization (BiDAVS) that can be highly parallelized. Our proposed BiDAVS considers position information of input sequence and effectively capture long-range temporal dependencies of sequential frames by computing bi-directional attention. Extensive experiments on two popular benchmark datasets, i.e., SumMe and TVSum, show that our proposed model outperforms state-of-the-art approaches.

Read the paper · More papers on PaperTik