Vision-To-Voice: An Intelligent Caption Generation and An Avatar-Guided Assistive System for The Visually Challenged

D. Karthika, S. Balamurugan · International Journal of Basic and Applied Sciences · 2025

Access to visual information is critical to independent living, but individuals who are visually ‎impaired continue to encounter obstacles in seeing and understanding their environment. ‎Conventional assistive technologies like screen readers and object detectors are generally weak ‎in semantics, contextuality, and interactivity, which makes them inadequate for actual use in ‎real-world settings. To address the shortcomings, an intelligent and interactive assistive system ‎, Vision-to-Voice, is architected to transform static visual information into meaningful verbal ‎descriptions with deep learning and real-time avatar-guided narration. The proposed system ‎presents a new end-to-end image captioning architecture that incorporates improved ‎preprocessing, a dual-stream deep feature extraction flow, and a context-aware caption ‎generation model. During preprocessing, images are normalized and denoised to ‎enhance feature clarity. The feature extraction is conducted through a hybrid ‎ResNet50+custom-convolutional stream architecture that combines global and local ‎representations from pre-trained ResNet50 and custom-trained convolution streams. A well-designed dataset of 1,600 visually diverse images and 8,000 respective human-written captions ‎is employed, with 80% reserved for training and 20% for testing. Five descriptions are assigned ‎to each image, promoting semantic diversity during training. The captioning model is trained to ‎learn from several contextual cues, making it possible to generate rich, human-like captions. ‎The performance of the system is measured quantitatively in terms of accuracy, precision, ‎recall, and F1-score, all of which show significant improvements over traditional single-stream ‎or template-based approaches.‎ To facilitate real-time use, the trained model is embedded in a graphical user interface (GUI) ‎with an intuitive design for simple navigation. The interface accommodates image loading, ‎captioning, and animated speech narration. A 2D avatar is also aligned with the synthesized ‎speech, visually realizing the captioned speech with audio-visual coherence throughout the ‎utterance. Captions are shown clearly in uppercase characters for improved readability. This ‎dynamic, multimodal feedback system enables a more inclusive and interactive experience for ‎visually impaired users. The system not only excels in generating captions with higher accuracy ‎but also provides a pragmatic and compassionate assistive solution with its harmony of cutting-edge vision-language modeling and human-centric design. User-oriented considerations and ‎test results all verify the framework's viability for real-world accessibility use cases, paving the ‎ground for future advancements in assistive AI‎.

Read the paper · More papers on PaperTik