Deep Neural Networks for Image Caption Generation

K. Tejaswi, Mohan Lal · International Journal of Research Publication and Reviews · 2025

This project explores the rapidly growing field of image caption generation, which merges the strengths of computer vision and natural language processing to automatically produce human-like textual descriptions of images.The proposed framework utilizes Convolutional Neural Networks (CNNs) for extracting deep visual features and Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units for generating fluent and contextually meaningful captions.Trained on large-scale benchmark datasets, the model effectively captures complex visual-semantic relationships, enabling accurate and grammatically coherent caption generation.Building upon this core model, an Advanced Image Caption Generator web application is developed, designed to deliver 4-5 optimized captions per image, while also offering advanced capabilities such as object, color, and scene detection.To increase accessibility and usability, the system supports real-time translation into more than 40 languages, with a particular focus on Indian languages such as Telugu and Hindi.Optimized for performance, the application processes each image in just 3-4 seconds, combining speed, accuracy, and a user-friendly interface to provide a seamless experience.Experimental results demonstrate that this approach achieves competitive performance when compared with existing state-of-the-art models, showing significant improvements in accuracy, fluency, and multilingual adaptability.This work highlights the potential of deep neural networks in real-world applications, including digital accessibility for the visually impaired, intelligent image retrieval, content management, and assistive technologies, while paving the way for future advancements in multimodal artificial intelligence systems.

Read the paper · More papers on PaperTik