Machine Learning Approach for Image Caption Generation
Subasish Mohapatra, Jyoti Ranjan Nayak, Samita Bala Routray, Devismita Sahoo, Subhadarshini Mohanty · 2026
This work aims to develop an advanced Image Caption Generator by integrating deep learning techniques from the available domains of NLP and computer vision. Utilizing the data set named Flickr8k effectively produces descriptive captions for the images by accurately interpreting visual content. The workflow encompasses meticulous preprocessing of images and captions, where images are encoded using a pre-trained ResNet50 model, and captions are vectorized to form structured data inputs. The innovative model architecture features separate processing branches for images and text, which are subsequently concatenated to yield a unified output. This paper highlights the synergistic potential of Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN) to construct a robust system capable of producing contextually relevant captions for images. Image captioning has significant applications across various sectors. By combining the strengths of CNNs in visual recognition with RNNs’ proficiency in language generation, this paper not only demonstrates the current capabilities of AI in understanding and describing visual content but also sets a foundation for future advancements in multimedia content processing and AI-driven user interfaces.