VIEW: Optimization of Image Captioning and Facial Recognition on Embedded Systems to Aid the Visually Impaired
Akshit Bhalla, S Goutham, Kevin Prakash, T Sanjana · 2021
Recent advancements in object recognition, face recognition, and optical character recognition allows for obtaining detailed descriptions of an environment. However, they present computational challenges when run on low-powered, embedded devices. The system described in this paper, Visual Interpreter of Environment Wizard (VIEW), combines these techniques in a computationally efficient manner designed to run on portable, low-powered devices. VIEW can assist people with visual impairment by providing natural language descriptions of their surroundings, information about the relative movement of nearby objects, and navigational utilities. For object recognition, we have implemented a modified deep neural network architecture that consists of purely convolutional layers with quantization and does not contain any fully connected layers to reduce the size of the model to provide optimal performance. The model is trained on the Common Objects in Context (COCO) dataset for multiple objects of various classes and processes the image during inference only once which enables faster image processing. All features of the system integrate voice commands for hands-free use. On a Raspberry Pi 4, the system achieves near real-time image captioning performance.