Searching a Video Database using Natural Language Queries

M Shubha, Kritika Kapoor, M Shrutiya, H. R. Mamatha · 2021 International Conference on Emerging Smart Computing and Informatics (ESCI) · 2021

This paper describes an application that achieves voice based natural language query, search and extracted video segment playing after the search in order to query the content of the videos in a user-friendly manner. Two different models were explored for the same. The first model is implemented using an image captioning approach. Two different image captioning methods are used for creating tracklets, namely Densecap and NeuralTalk2. NeuralTalk2 generates a single appropriate caption for the entire image whereas Densecap generates multiple captions corresponding to specific regions of interest in the image. These captions are used to preprocess the video and create semantically similar tracklets. Given a video and a voice based natural language query, this system will produce video tracklets from the video that are semantically relevant to the query. The second model uses an audio processing approach. Here, first the transcripts generated by YouTube are collected. The voice query is taken as input and the most relevant segments of the video are retrieved using on-the-fly generation of tracks and merging if required. For finding Semantic similarity in both the models, first Universal Sentence Encoder (by Google) which uses a deep averaging network encoder (DAN) for converting the sentences into 512 dimensional vectors is used and then cosine similarity between the vectors is calculated.

Read the paper · More papers on PaperTik