Scene Description Using Keyframe Extraction and Image Captioning
Daniel Lester Saldanha, Raghav T Kesari, K Rahul Srinivas, Senthil Kumaran Vijayalakshmi Natarajan · 2023
Multimedia content has evolved drastically over the past few years and has necessitated study into audio-video content. Although visual media is widely available, the accessibility for individuals with visual impairments is inadequate. Most successful attempts at improving scene understanding require laborious manual work as automation is minimal. In recent years, there has been significant research in the field of generating descriptions for videos and the processing of images. Our primary goal is to generate accurate descriptions of the scenes in a given video and to present a framework for automating the generation of these descriptions. Through our framework, we were able to obtain better METEOR scores than previous automated implementations with a METEOR score of 0.46 with the MSVD dataset and 0.134 with the MPII Movie Description Dataset. The implementation divides the process into many incremental steps to generate descriptions for all the individual scenes in a given video. The framework involves preprocessing the given video, followed by a process known as keyframe extraction. This generates an output of the most relevant frames in the given video. This is fed as input to the image captioning algorithm that generates a caption for every single keyframe. Following this, we use a summarisation method to obtain a description of every scene in a given video. Our framework provides the ability to separate the task into many steps or modules, each of which can be separately improved. This allows the framework the scope to constantly evolve as more research and breakthroughs are achieved in the relevant fields.