Intelligent Image Captioning with InceptionV3, LSTM, and PySpark Integration

Thanu Kurian, Jimsha K Mathew, Syam Dev R S, P K Anand, A Vignesh, B S Rakshith · 2025

Image captioning is a challenging task in artificial intelligence that involves generating descriptive captions for images automatically. In this project, we propose a novel approach leveraging advanced technologies such as PySpark, LSTM, and InceptionV3 to develop an effective image captioning system. We harness the power of PySpark, a distributed computing framework, to efficiently process large-scale image datasets and extract high-level features from images using the InceptionV3 convolutional neural network (CNN) model. These features capture the semantic information present in the images and serve as input to the caption generation model. The caption generation model utilizes LSTM neural network as the decoder component. LSTM is well-suited for sequential data processing and is capable of generating coherent and contextually relevant captions based on the extracted image features. The InceptionV3 model acts as the encoder, extracting meaningful visual features from input images, while the LSTM decoder generates captions by decoding these features into natural language descriptions. This multimodal approach enables the model to understand and describe the content of images accurately. Through experimentation and evaluation on diverse image datasets, our proposed system demonstrates promising results in generating accurate and human-like captions for a wide range of images. By integrating PySpark, LSTM, and InceptionV3, our image captioning system achieves state-of-the-art performance, highlighting the effectiveness of leveraging advanced technologies in solving complex AI tasks.

Read the paper · More papers on PaperTik