Integrating Vision Transformers and Text-to-Speech System for Converting Object Detection Outputs to Audio Descriptions using AI

M.C. Babu, Saravanan V, K. Ramanan · 2025

This paper presents an integrated system that combines Vision Transformers (ViTs) for object detection with Text-to-Speech technology to create a real-time assistive tool for the visually impaired. The system exploits DETR to detect objects in a scene with very high accuracy and contextual depth, while gTTS provides natural, multilingual auditory feedback. This well-trained Vision Transformer model serves images and delivers a very detailed description of objects, such as spatial relations, which can be converted into speech in any language user prefers. Some results of evaluations from COCO show the working of the system: it achieves 90.3% for mean Average Precision (mAP) and 4.5 as Mean Opinion Score (MOS) for the speech quality. It works in real time, and on average, processes frames in 0.47 seconds, with latency of 0.8 seconds from detection to generation of feedback. Although the system is found wanting in low-light conditions and offline usability, its strengths in accuracy, accessibility, and multilingual support mark it as a transformative assistive technology.

Read the paper · More papers on PaperTik