Neural Network-Based Image Captioning
Samatha R Swamy, Shashank S Dengi, Subham Kumar Gupta, Sourav Nayak, Tarun Gowda H D · 2025
Image captioning integrates computer vision and natural language processing to enable AI to generate descriptive text for visual content. This approach combines Convolutional Neural Networks (CNNs) to extract image features with Long Short-Term Memory (LSTM) networks to produce fluent captions. Built using PyTorch, the system is trained on the COCO dataset, a standard for image captioning research. The process involves three key steps: preprocessing the COCO dataset to create caption embeddings, employing a pre-trained ResNet-50 model to capture visual features, and training an LSTM to generate captions sequentially. This multimodal framework tackles challenges like managing varying caption lengths, aligning visual and textual elements accurately, and ensuring natural language output. By merging CNNs and LSTMs, the system balances visual analysis with text generation. It is designed for adaptability, with future enhancements like attention mechanisms or transformer models planned to enhance contextual precision. Applications include aiding visually impaired individuals with detailed image descriptions, improving search functionalities, and automating content labeling for datasets. Future developments will focus on handling complex or ambiguous images and expanding the model to support diverse languages and specialized domains, increasing its utility in assistive technologies, social media management, and data annotation.