Auto-Image Caption Generator using Artificial Intelligence

Abhyuday Srivastava, Saransh Singh, K.C. Sriharipriya, P. Sundresan, G.N. Sumathi · 2025

This project presents an automated image captioning system that integrates deep learning techniques from computer vision and natural language processing. The architecture combines a pre-trained Convolutional Neural Network (CNN), specifically VGG16, for visual feature extraction, with a Long Short-Term Memory (LSTM) network for sequential text generation. Transfer learning is used to convert images into rich feature vectors, while tokenized captions guide the language model during training. An attention mechanism is incorporated to dynamically focus on salient regions of the image, significantly improving caption accuracy and contextual relevance. Practical applications include accessibility support for visually impaired users, automated content tagging, and intelligent image indexing in domains such as media and surveillance. The system is deployed via Streamlit, offering a scalable and user-friendly interface for real-world deployment.

Read the paper · More papers on PaperTik