Enhancing Accessibility for Visually Impaired Users: A BLIP2-Powered Image Description System in Tamil
Smrithika Antonette Roslyn, A Negha, Vishnu Sekhar R · 2024
Image captioning is the task of analyzing an image and expressing its content using natural language. It involves computer vision, image analytics, object recognition and natural language processing. This field has been attracting greater popularity recently resulting from its possible applications in numerous domains, including accessibility for those with visual impairments, which lacks in the Tamil community due to limited availability of reliable and accurate tools for creating context-aware image captions in Tamil. The primary objective of this research is to create a context-aware image captioning system capable of giving a thorough explanation of an image in Tamil to assist in the lives of the visually impaired. The proposed technique will leverage Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (BLIP-2) technology and generate image descriptions in Tamil and convert them to audio format. This will extract visual features using a frozen image encoder, and process them using a large language model to produce natural language descriptions in Tamil. This research is based on BLIP's ability to process linguistic as well as visual data efficiently along with being able to provide accurate, natural-sounding image descriptions in the target language, making them a best fit for the task.