Virtual Assistance System for Visually Impaired people using Deep Learning

Faiyaz Ahmad, Mohd Zeeshan Ansari, Mohd Tayyab, Sharique Shuja, Muhammad Zaki · Procedia Computer Science · 2025

Assistance is an important necessity for visually impaired people, as it helps them to navigate in challenging conditions in day-to-day life. Object detection, distance and position estimation of nearby objects are crucial tasks to user navigation, which helps to navigate them in outdoor environments. Consequently, utilizing the knowledge of the surroundings by the caption generation of view is also essential to assistance. This paper presents a novel framework that includes object detection, distance and position estimation to generate a generic description of the view. The framework uses best in class pre-trained Yolov5 for object detection, whose inference speed is better than its predecessor, subsequently, a custom model trained on a Kitty dataset for distance estimation. Attention-based model for caption generation, which uses Efficient-Net architecture as the backbone for feature extraction. It is observed that the implemented distance model provides best results in 3 out of 4 parameters, while the proposed Caption Generation framework provides best BLEU scores among multiple models explored. Yolov5 provides 42.5 mAP on COCO dataset. For the Kitty dataset Custom distance estimation model with 296000 parameters, a test RMS Log of .0232 and .1627 squared relative error is obtained.

Read the paper · More papers on PaperTik