Comparative study of different LLM's for Captioning Images to Help Blind People

Jyoti Madake, Mrunal Sondur, Sumit Morey, Atharva Naik, Shripad S. Bhatlawande · 2024

The proposed system leverages LLMs, including GPT2, DistillGPT2, BERT, and RoBERTa, to provide detailed scene descriptions for the visually impaired. It employs an Encoder-Decoder architecture, with the Vision Transformer as the encoder and a distilled GPT-2 model as the decoder, facilitating the generation of comprehensive image captions. Training used a diverse dataset of around a hundred thousand samples on hardware equipped with an Nvidia Tesla V100 PCIE card and two Intel Xeon Silver CPUs. The model achieved a ROUGE score of 21.69%, signifying its potential to generate captions closely resembling human descriptions after training on suitable LLMs in the near future. This research has profound implications for enhancing the independence, education, and employment prospects of the visually impaired.

Read the paper · More papers on PaperTik