Shot Level Egocentric Video Co-summarization
Abhimanyu Sahu, Ananda Shankar Chowdhury · 2018
Video co-summarization has emerged as an important problem in the areas of computer vision and multimedia communities. In this paper, we present a novel approach of co-summarizing egocentric videos at shot level. Our solution pipeline consists of three major components. We develop a new way of characterizing egocentric video frames by computing the differences in contrast, entropy and optic flow values between a central region and the surrounding region in a frame. This is termed as the center-surround model. Visual similarity between a test video shot and a database video shot is next derived using a game-theoretic framework. Each video shot is modelled as a player and the expected pay-off difference between any two such players at mixed Nash equilibrium is deemed as the similarity between them. A weighted bipartite graph is constructed next between the shots in a test and in a database video. Game-theoretic similarity values are deemed as the weights. Maximum Cardinality Minimum Weight matching in the bipartite graph yields non-greedy shot correspondences. Best matched shots from the test video are used to form the summary. Experimental comparisons on standard datasets clearly indicate the advantage of our solution.