Audio Description of Images using Deep Learning
A.Pramod Kumar, Vennam Vignesh, K R Abhiram, Yegireddy Chakradeep, Mucha Srinikesh · 2024
Human brains are capable of making sense of the information from the surroundings with visual cognition input most of the time. People with visual impairment find it greatly challenging to interact with their environment and if an audio description of their environment is available, it significantly increases their accessibility. This study is interested in generating audio description of images using deep learning architectures. As a general principle, most existing methods for extracting information from images apply CNNs (Convolutional Neural Networks), which are deep learning architectures designed for image classification tasks. This study discusses a model that addresses the necessity of taking advantage of the natural language processing technologies such as transformers. Being the most common thing in Natural Language Processing (NLP), they are used extensively for image-to-text extraction, therefore, they have better descriptions. This study has added the aspect of audio description to the system which can be used as the practical application of the model in visual aids for better understanding of the environment.