Vision-To-Voice: An Intelligent Caption Generation and An Avatar-Guided Assistive System for The Visually Challenged
D. Karthika, S. Balamurugan · International Journal of Basic and Applied Sciences · 2025
Access to visual information is critical to independent living, but individuals who are visually impaired continue to encounter obstacles in seeing and understanding their environment. Conventional assistive technologies like screen readers and object detectors are generally weak in semantics, contextuality, and interactivity, which makes them inadequate for actual use in real-world settings. To address the shortcomings, an intelligent and interactive assistive system , Vision-to-Voice, is architected to transform static visual information into meaningful verbal descriptions with deep learning and real-time avatar-guided narration. The proposed system presents a new end-to-end image captioning architecture that incorporates improved preprocessing, a dual-stream deep feature extraction flow, and a context-aware caption generation model. During preprocessing, images are normalized and denoised to enhance feature clarity. The feature extraction is conducted through a hybrid ResNet50+custom-convolutional stream architecture that combines global and local representations from pre-trained ResNet50 and custom-trained convolution streams. A well-designed dataset of 1,600 visually diverse images and 8,000 respective human-written captions is employed, with 80% reserved for training and 20% for testing. Five descriptions are assigned to each image, promoting semantic diversity during training. The captioning model is trained to learn from several contextual cues, making it possible to generate rich, human-like captions. The performance of the system is measured quantitatively in terms of accuracy, precision, recall, and F1-score, all of which show significant improvements over traditional single-stream or template-based approaches. To facilitate real-time use, the trained model is embedded in a graphical user interface (GUI) with an intuitive design for simple navigation. The interface accommodates image loading, captioning, and animated speech narration. A 2D avatar is also aligned with the synthesized speech, visually realizing the captioned speech with audio-visual coherence throughout the utterance. Captions are shown clearly in uppercase characters for improved readability. This dynamic, multimodal feedback system enables a more inclusive and interactive experience for visually impaired users. The system not only excels in generating captions with higher accuracy but also provides a pragmatic and compassionate assistive solution with its harmony of cutting-edge vision-language modeling and human-centric design. User-oriented considerations and test results all verify the framework's viability for real-world accessibility use cases, paving the ground for future advancements in assistive AI.