Enhancing Image Captioning through the Integration of Image Processing and NLP
Sakshi Ugale, Druwil Jain, Diptee Vishwanath Chikmurge, Sunita Barve · 2024
The study offers a novel method for captioning images by combining Gated Recurrent Units (GRU) and Convolutional Neural Networks (CNN). A feature extractor called VGG19 can be used to take complex photos and extract high-level visual representations from them. The aforementioned visual inputs are subsequently fed into a language model that employs a GRU and Bi-directional LSTM, two widely recognised models for handling sequential dependencies, to produce captions that make sense and are contextually relevant. Selective information update is made possible by the GRU’s gating mechanism, making it simpler to capture long-range dependencies in the output captions. The combination of VGG19, GRU, and Bi-directional LSTM provides a comprehensive and hierarchical method for photo captioning that extracts rich visual information required for comprehending complex picture content. GRU’s gating technique makes it easier to analyse specific information, which improves the model’s comprehension of complex relationships found in output captions. In order to provide more accurate image descriptions, the suggested methodology seeks to improve the synergy between visual and verbal knowledge. Comparing experimental assessments against modern technique approaches on benchmark datasets with metrics like BiLingual Evaluation Understudy (BLEU), METEOR, and CIDEr shows encouraging results. In addition to its potential uses in computer vision(CV), accessibility, and human-computer interaction, this work advances the field of image caption development.