Vision Transformer Based Vision Enhancement for Visually Impaired Individuals
Mohammed Ovaiz A, A Yogaraj, Rani K.S., Madhan Kumar V, Mohammed Kabir M, K Veeramuthu · 2024
The goal of visual implants is to create artificial vision that can partially restore function. It can enhance the quality of life for visually challenged individuals by allowing them to feel light, even after years of darkness, by the use of 60 microelectrodes implanted in the retina. The artificial vision that is made possible by current visual system stimulators has very poor resolution because of their small number of microelectrodes. Numerous researchers have sought to enhance artificial vision produced by low-resolution implants through the application of machine learning and image processing techniques. Because phosphine pictures have low resolution, users report unhappiness with the Retinal Prosthesis System. This underscores the important need for targeted research aimed at improving visual clarity and user pleasure in general. This research proposes simulating artificial vision in which the visually impaired user receives information synthesized by the system through a low-resolution photo courtesy of a visual implant. Through the use of Vision Transformer, the technique gathers useful data about people in the immediate vicinity of the visually impaired person, including their number, familiarity, gender, approximated ages, facial emotions, nearby items, and approximate distances. The information obtained from the user's glasses' camera frames is used to create signals that are then sent into a visual stimulator, offering a potentially effective way to improve the visual experience for those who are visually impaired. In order to facilitate economical real-time implementations in an independent portable system, an algorithm that best suits each feature is chosen based on its accuracy and time complexity. The proposed approach uses audio to provide crucial information about those in close proximity to a visually impaired person, enabling them to converse with others more comfortably. This paper can thus be taken into consideration for some next-generation visual implant systems. The accuracy of the CNN is 93% and Specificity comparison of CNN is 92% respectively.