Blind Leap Real-Time Object Recognition with results converted to Audio for Blind People
R Jaya · Zenodo (CERN European Organization for Nuclear Research) · 2019
This project tries to change the visual world into the audio world. It has the likelihood to inform blind people about the objects as well as their spatial locations. The objects that are detected at the scene are represented by their names and are then transformed to speech. Their spatial locations are encoded into the 2-channel audio with the help of 3D binaural sound simulation. The system is collected of various modules. The video is captured by a portable camera device (Raspberry Pi with Noir Camera) on the client side. It is then streamed to the server for real-time Object recognition with existing object detection models (YOLO). The 3D location of the objects is determined by the location and the size of the bounding boxes using the detection algorithm. A 3D sound generation application, built on Unity game engine then renders the binaural sound keeping the locations encoded. The transmission of the sound to the user happens with Bluetooth/3.5 jack earphones. The sound is played at an interval of a few seconds or when the recognized object differs from the last one - depends which one is the earliest.