A Survey on Generating Audio with Captions for a Live Video using Neural Networks
Archana Uriti, Durga Prasad Kakarla, Deepika Jinugu, Kanaka Nihitha Silla, K.T. V.N.S. Sai Krishna, Chunduru Anilkumar · 2023
Live video caption generation is a challenging task where the generated output should be accurate generalized captions for a video content and also providing audio will also get more attention. It is used in various applications like surveillance, military operations etc. Generating captions for a real time video is very difficult where the caption generator module has to semantically understand video content and translate it into meaningful captions to make the visually impaired people understand. As there are various approaches available to implement the live video captioning, this research study discusses about the various methodologies used for live video captioning and also works to overcome the existing challenges. This research study analyses the complete state-of-art models, datasets and evaluation metrics used for video captioning in the existing models. The main goal of this research study is to provide a clear understanding on the already existed works and the future scope.