Image Caption Generation for the Visually Impaired Using Deep Learning

Shruti Deshpande · International Journal for Research in Applied Science and Engineering Technology · 2025

Abstract: Providing written descriptions of visual content, image caption generation has become a vital assistive technique for people with vision impairments. In this study, an improved deep learning framework is presented to produce precise and contextually rich image captions intended for assistive technology applications. Our suggested architecture, which we refer to as ViT-BiLSTM-Attention (VBLA), combines a Vision Transformer (ViT) encoder with a bidirectional LSTM decoder enhanced by an attention mechanism. We tested our model on a novel dataset that has been specially selected for assistive technology applications, as well as on common datasets like Flickr30k and MS COCO. With a BLEU-4 score of 0.382, METEOR score of 0.417, and CIDER score of 1.142, the experimental results show that our method outperforms current approaches and achieves state-of-the-art performance. Perform a thorough user research with visually challenged volunteers to assess our approach’s practical efficacy. In this work, special difficulties of developing image captioning systems for assistive technology are addressed. These difficulties include the need for detailed spatial descriptions, the recognition of important objects, and natural language generation that gives users with visual impairments priority over irrelevant information.

Read the paper · More papers on PaperTik