Deep Learning-Based Image Captioning: A Hybrid CNN-LSTM Approach.
V. Pravallika, Vadduri Uday Kiran, B. Rahul, Nizampatnam Neelima, G. Rishi Patnaik, DR. Sreejyothshna Ankam · International Journal of Research Publication and Reviews · 2025
In today's digital age, image captioning has become a crucial tool for bridging the gap between visual content and natural language.This project aims to develop an image caption generation model that automatically produces descriptive and coherent text for a given image.Image captioning plays a significant role in computer vision and natural language processing, with applications in platforms like Facebook and Google Photos for image segmentation and organization.Additionally, it can automate tasks that require human interpretation of images, making it valuable for accessibility and content management.To achieve this, the project utilizes deep learning techniques, specifically Convolutional Neural Networks (CNN) for visual feature extraction and Long Short-Term Memory (LSTM) networks for sequential text generation.The model is trained on the Flickr8k dataset, which contains 8,000 images, each paired with five captions.This approach enhances the accuracy and relevance of generated captions, making it useful for various real-world applications in accessibility, search engines, and multimedia platforms.