Attention Based CNN-RNN Hybrid Model for Image Captioning

Anisha Raphael, S Abisri, E Anitha, S Ritika, Manju Venugopalan · 2024

Image captioning enables computers understand and interpret visual content by bridging the gap between natural language processing and computer vision. The model proposed in this paper for image captioning that employs a hybrid deep learning method and includes an attention mechanism. Combining the strengths of convolutional neural networks (CNNs) for image interpretation and recurrent neural networks (RNNs) for caption generation allows machines to provide more accurate and contextually appropriate descriptions. The CNN+RNN with attention technique improves captioning by dynamically focusing on relevant image regions, and the CNN+RNN encoder-decoder structure serves as a robust foundation for textual descriptions. In contrast, the Vision Transformer (ViT) with GPU acceleration provides greater feature extraction and faster processing, using global context for potentially more accurate and efficient caption production. This work presents a detailed approach to image captioning using a well-chosen dataset from Kaggle. This dataset features images with human-written captions, which helps train the model more effectively. The models experimented were CNN+RNN with attention mechanism, CNN+RNN encoder-decoder and ViT+GPU2. The best results are reported from DenseNet201 which is based on CNN+RNN with attention mechanism achieving an average BLEU score of 0.668 and an average ROUGE-L score of 0.746.

Read the paper · More papers on PaperTik