Visual content based video retrieval on natural language queries

Ravi Bansal, Sandip Chakraborty · 2019

With the emergence of video-sharing websites uploading thousands of videos per day, content based video retrieval has become an important research challenge. Existing approaches for content based video retrieval lag robustness and efficiency in terms of the dynamics and variations of videos and queries expressed in natural language. In this paper, we develop a novel approach for bi-directional retrieval of videos and its descriptions, represented through sentences in natural language, by combining various state-of-the-art learning and video/sentence representation techniques. We first develop a base model which is based on sentence embedding of captions generated from video content. Further we extend the base model with a joint learning of latent embedding space model, where we use a two-branch neural network to learn the shared embedding space of videos and corresponding sentences expressing its description. The network is trained using a maximum-margin loss and uses neighborhood information for structure-preserving constraints in joint embedding space. Our experiments with three different datasets show promising results for video and sentence retrieval in comparison with various baseline models.

Read the paper · More papers on PaperTik