Synergistic Fusion of CNN and BiLSTM Models for Enhanced Video Captioning

Pankaj Sharma, Sanjiv Sharma · 2024

In the realm of video captioning systems, the significance of leveraging CNN models for feature extraction cannot be overstated. These deep features, extracted from video frames, are seamlessly integrated with textual data features through a probabilistic matching approach, establishing crucial connections with the target text data. This study delves into the effectiveness of utilizing pretrained standard convolutional neural networks for feature extraction from video frames, with a keen focus on identifying the most suitable model to bolster video captioning systems. Through an exhaustive analysis, we explore various combinations of recurrent neural networks (RNN) with LSTM and GRU layers, coupled with the integration of Bidirectional layers, aiming to pinpoint nuanced model requirements for optimizing BLEU scores. Evaluation on the MSR- VTT dataset showcases notable enhancements, particularly when employing the MobileNet-V3 model for video frame feature extraction in conjunction with a BiLSTM-based model for subsequent text data processing. These findings culminate in substantial improvements in captioning outcomes, underscoring the potential of this integrated approach.

Read the paper · More papers on PaperTik