Pixels to Phrases: Bridging the Gap with Computationally Effective Deep Learning models in Image Captioning

K Serath Chandra, R. Poda, Ardra Vinod, R. Arun · 2023

Image captioning combines computer vision and natural language processing using deep learning for generating descriptive image captions, transforming how we perceive and interact with visual data, enhancing comprehension, and enabling multimodal communication. This Research aims to use lightweight deep-learning models for image captioning, bridging the gap between visual understanding and caption generation. Models adopt an encoder-decoder architecture, using CNNs (VGG16, InceptionV3, DenseNet, EffcientNet, MobileNet) as encoders for extracting high-level visual features. Complementing the encoder, a recurrent neural network (RNN) such as Long-Short-Term-Memory (LSTM) serves as the decoder, generating captions word by word. The transfer learning technique is incorporated to enhance the image caption generator’s performance. This involves pretraining all the CNN models with the ImageNet dataset, followed by fine-tuning using the Flickr8k dataset specifically for caption generation. By leveraging this approach, the proposed system achieves improved captioning quality and accuracy, showcasing the significance of transfer learning in the domain of image caption generation. In addition to enhancing caption quality, this research work also focuses on reducing the caption generation time and computational complexity of the model. The proposed image generator with MobileNet and LSTM outperformed when compared with other models in all aspects with a better bleu Score of 55.15% and lesser computational complexity of 5.09 million trainable parameters.

Read the paper · More papers on PaperTik