Integrating Convolutional and Recurrent Networks for Image Caption Generation: A Unified Approach

I Nandhini, L. Prasanth, T Nagalakshmi, D. Manjula · 2024

Within the realm of natural language processing and computer vision, the synergy between Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) systems has emerged as some powerful paradigm for image captioning. This paper presents a comprehensive solution for visual and temporal understanding in the image captioning domain by presenting a unified framework that efficiently blends the strengths of CNN and LSTM architectures. Our proposed model takes advantage of the spatial hierarchies captured by CNNs in images and the sequential dependencies modeled by LSTMs in natural language. The CNN component efficiently retrieves high-degree features to result in coherent and contextually rich image captions. The hybridization of both of these networks creates a symbiotic relationship, addressing the challenges associated with both spatial and temporal aspects of image interpretation. We carried out in-depth tests on benchmark datasets, demonstrating the superiority of our hybrid CNN-LSTM architecture over standalone models. Our results showcase improved captioning exactness, fluency, and contextual relevance. Our proposed model achieved highest BLEU score $\mathbf{0. 8 1}$ when compared with existing methods. In addition, we investigated the interpretability of our model, shedding light on the visual and temporal cues learned by the hybrid architecture.

Read the paper · More papers on PaperTik