Enhancing Video Captioning: Harnessing CNN Model Features with BiLSTM Integration
Aditee Abhijeet Salokhe, Hemant Appa Tirmare, Hridaynath Pandurang Khandagale · 2024
In video captioning systems, the pivotal role of feature extraction using CNN models is underscored. The deep features derived from these models are subsequently integrated with text data features through a probabilistic matching approach to establish connections with the target text data. This paper explores the efficacy of utilizing pre-trained standard convolutional neural networks for extracting features from video frames, aiming to identify the most suitable model for enhancing video captioning systems. Through a comprehensive analysis, the paper examines the combination of recurrent neural networks (RNN) with LSTM and GRU layers, along with the incorporation of Bidirectional layers, to determine nuanced model requirements for achieving the highest BLEU score. Performance evaluation on the MSR-VTT dataset reveals enhanced results when employing the ResNet-101 model for video frame feature extraction and a BiLSTM-based model for subsequent text data processing, ultimately contributing to improved captioning outcomes.