Visionary Narratives: Unveiling the Potential of an Automated Image Caption Generator through Deep Learning and Multidimensional Analysis

Kumar Keshamoni, Pabbala Priyanka · 2024

This paper introduces an image caption generator, implemented from Yumi's Blog, which employs a deep learning model trained on the FLICKR_8K dataset. The generator leverages computer vision and natural language processing to produce descriptive captions for images. Potential applications range from aiding visually impaired individuals to medical and geospatial image analysis. The project encompasses a comprehensive workflow; including data cleaning, feature extraction using VGG-16, LSTM model building, and evaluation through BLEU scores. The model's performance is showcased through use cases, demonstrating its potential impact on diverse fields such as accessibility, advertising, and healthcare. The study concludes with insights into hyperparameter tuning and the acknowledgment that while the model yields promising results, further refinement is possible for enhanced caption generation.

Read the paper · More papers on PaperTik