Blind Assistance: Object Detection with Voice Feedback

Mosarrat Shazia Kabir, Syeda Karishma Naaz, Md. Tahmid Kabir, Md. Shahriar Hussain · 2023

In the contemporary world, the field of computer vision is experiencing rapid growth. Computer vision involves enabling a computer to attain a comprehensive comprehension of digital images. When individuals observe objects, their neurons facilitate swift reactions within milliseconds. However, this doesn't hold true for those with visual impairments. Negotiating their surroundings and responding to objects isn't as effortless. According to the World Health Organization (WHO), approximately 39 million individuals globally suffer from significant vision loss [1]. The proposed Blind Assistance object detection with voice feedback is being designed to assist visually impaired individuals. This system has been trained utilizing the You Only Look Once (YOLO) algorithm and the Single Shot Detection (SSD) Mobilenet. The models employed in this project include SSD Mobilenet v2, YOLO v4, and YOLO v7. The Microsoft COCO dataset, featuring the capability to classify more than 80 object categories, serves as the foundation for this initiative. Moreover, a counting function has been integrated to enumerate object classes in real-time, providing a comprehensive count and identification of objects. This information is then converted into voice feedback using the Google Text to Speech API through the gTTS package. This voice feedback empowers the visually impaired to perceive the visual world that would otherwise be inaccessible to them.

Read the paper · More papers on PaperTik