Image Captioning-A Comprehensive Encoder-Decoder Approach on Flickr8K
Mithilesh Mahajan, Saanvi Bhat, Madhav Agrawal, Gitanjali R. Shinde, Rutuja Rajendra Patil, Gagandeep Kaur · 2025
This paper proposes a new encoder-decoder frame-work based on Convolutional Neural Networks (CNN) for feature extraction and Long Short-Term Memory (LSTM) networks for image caption generation. Image captioning, the process of generating descriptive textual content for images, has diverse applications, including accessibility tools for the visually impaired, virtual assistants, social media automation, and satellite imagery analysis in critical domains like defense and disaster management. The proposed model will leverage the power of CNNs for extracting high-level image features and LSTMs for generating semantically coherent captions. Rigorously evaluated on the Flickr8k dataset, the model achieved competitive BLEU scores, showing its efficacy in generating accurate and meaningful descriptions. The methodology integrates advanced preprocessing techniques and efficient training strategies, ensuring robustness and scalability. This re-search focuses on the need to use state-of-the-art machine learning techniques in the solving of real challenges, and some of the potential routes to improving image captioning systems are suggested to include multilingual capabilities and integration with larger datasets. Contributions include a clear explanation of the encoder-decoder approach, detailed metrics for evaluation, and insights into correcting limitations including dataset size and complex scene representation. This work is a step toward further progress in automated image-to-text systems, integrating precision, relevance, and adaptability.