Vision to Voice: Transforming Images into Audio Descriptions with Deep Learning

N. Sunanda, Vodnala Sharan Adhvy, Kanvapuri Sai Praneetha, Katikireddy Sai Sahithi, Mohammed Shaji Affan · 2025

The visually impaired have great difficulties sensing and engaging with the visual environment. Activities such as exploring public areas, recognizing objects, or retrieving visual information prove to be challenging when not properly supported. Although assistive technologies are available, the majority are short on contextual awareness and natural interaction. In this paper, we introduce "Vision to Voice: Translating Images to Audio Descriptions with Deep Learning", a system that utilizes the BLIP (Bootstrapped Language-Image Pre-training) model to produce precise, context-aware image captions, which are synthesized into speech via text-to-speech (TTS) synthesis. People can upload or photograph to receive real-time audio descriptions, making it easier with a more natural, voice-based interface. Our approach integrates progress in computer vision, natural language understanding, and deep learning to deliver semantically accurate and logically consistent descriptions. It could be applied in mobile apps, voice assistants, and blind-friendly self-navigating devices with offline availability, multilingual support, and AI-powered personalization.

Read the paper · More papers on PaperTik