Visionary Narratives: Unveiling the Potential of an Automated Image Caption Generator through Deep Learning and Multidimensional Analysis
Kumar Keshamoni, Pabbala Priyanka · 2024
This paper introduces an image caption generator, implemented from Yumi's Blog, which employs a deep learning model trained on the FLICKR_8K dataset. The generator leverages computer vision and natural language processing to produce descriptive captions for images. Potential applications range from aiding visually impaired individuals to medical and geospatial image analysis. The project encompasses a comprehensive workflow; including data cleaning, feature extraction using VGG-16, LSTM model building, and evaluation through BLEU scores. The model's performance is showcased through use cases, demonstrating its potential impact on diverse fields such as accessibility, advertising, and healthcare. The study concludes with insights into hyperparameter tuning and the acknowledgment that while the model yields promising results, further refinement is possible for enhanced caption generation.