Voice-based Image Captioning using Florence-2

Devanshi Mittal, Prashant Ahlawat, Tushar Nag, Om Brahma Shrimaroj · 2024

As visual information becomes the predominant tool for passing information in the society, people with visual impairment have a difficult time in comprehending images. This paper aims to meet this challenge by proposing a machine learning enabled method of producing descriptive text captions for images in a bid to provide a link to comprehend the visual content. This solution will involve processing and translating visual information into meaningful textual accounts in order to offer the visually impaired a way of gaining better understand and a more effective way of participating in the visual environment. The first strategy used a conventional image recognition technique, while the second strategy improved the method to improve captioning accuracy. This process was done not with the help of the old dataset but with the help of a new one in order to get improved results. Future work will concentrate on enhancing the system for special application of real-time captioning so that it can be easily used in real life.

Read the paper · More papers on PaperTik