Comparative Analysis on Generation of Image Captions Based on Deep Learning Models

G S Dakshnakumar, T. Jemima Jebaseeli · 2024

In the realm of image captioning, the quality of generated captions is paramount for effective communication of visual content. This analysis delves into deep learning to address the challenges faced by individuals with visual impairments, aiming to enhance their visual perception through innovative technologies. Traditionally, the visually impaired have relied on manual assistance and adaptive aids for navigation and understanding visual content. With the advent of deep learning, there is a unique opportunity to revolutionize this landscape. The study focuses on a comprehensive analysis of four distinct deep learning models: ResNet50, EfficientNetB7, DenseNet101, and VGG19. These models are rigorously trained for the task of image captioning, providing descriptive narratives for a diverse set of images. Results indicate that among the models, DenseNet101 and ResNet50 exhibit outstanding performance, achieving Bleu scores of 55.6 and 56.2, respectively. This astounding accomplishment shows how well deep learning models work at producing captions that nearly match reference captions. The study goes beyond numerical scores, delving into the nuanced characteristics of each model's architecture and its impact on captioning quality. This study contributes to the present developments in the field of assistive technology by demonstrating the possibilities of algorithms based on deep learning in supporting a more inclusive and autonomous experience for those with visual impairments. The knowledge gathered from this research opens the door for further developments in the use of AI to close the accessibility gap and empower the visually impaired community with enhanced tools for understanding and interpreting visual information.

Read the paper · More papers on PaperTik