Leveraging LSTM and CNN for Video Understanding
Kapuluri Nishitha, Bharathi Mohan G, Srinath Doss · 2024
This research introduces a sophisticated framework for real-time video captioning that seamlessly integrates a cohe-sive ensemble of machine learning models. The development of open-domain video descriptions that can manage variable-length input (sequences of words) and output (sequences offrames)while being sensitive to temporal structure is required due tothe complex dynamics present in real-world films. Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks are used to extract descriptive narratives and brief summaries from movies. LSTMs have demonstrated state-of- the-art performance in the generation of picture captions andare great at capturing sequential information. Our algorithm is trained on pairs of videos and matching texts to provide relevant captions for video events. It gains the ability to connecta string of words to a sequence of video frames. The modelnot only learns the temporal structure inside the sequence of frames, but also develops a language model for generating coherent and contextually relevant sentences. Our approach uses the VGG16 model to extract visual features, which results in a rich representation of the video information. These visual cues, along with the LSTM-based sequence-to-sequence model, enable the production of elaborate captions that faithfully capture the temporal dynamics and visual aspects of the video clips.