Object Detection and Video Analyser for the Visually Impaired
A Spoorthi Alva, R Nayana, Noorain Raza, Gambhire Swati Sampatrao, Koduru Bharath Subba Reddy · 2023
The proposed paradigm in this research can aid visually impaired individuals in developing a sense of their environment. It might be difficult to both see and feel what is going on around you if you have vision impairment of any kind. It’s challenging to move around and carry out tasks on one’s own. Studies estimate that there were more than 285 million visually impaired people in the world in 2012. The majority of the things in and around us share the same shape and size, which makes it much harder for someone who is visually impaired to grasp his surroundings. More than twice as many visually impaired people as the general population suffer from clinical depression. Understanding user accessibility and the kind of output that would work best in this situation proved to be a challenge. The delay has to be taken into consideration because the output had to be in real-time. The user may not benefit much from straightforward object recognition. So, using object recognition, the features have been investigated. A number of changes to different YOLO versions were made in order to develop object recognition. To recognize objects, this paper suggests YOLO V5. Giving a description of the video being fed into the model is the second feature. For this encoder-decoder sequence has been put into use. The Model has been trained and examined using a variety of data sets, all of which have been further described. The next crucial element is distance estimation, which detects things in front of the user and alerts them if they are approaching too closely. In answer to all these demands, the proposed model with three attributes can benefit the blind and be supplied as a solution. The model processes the video or photo input to identify the things seen in it. LSTM (Long Short-Term Memory) and VGG (Visual Geometry Group)16 are used to characterize the video. Triangular similarity and frozen graphs are used to measure distance and inform users if a subject is approaching the camera too closely. The user is provided with an audio description of the video being used as input. Both tools deliver audio as their output.