Assistant for people with visual disabilities through the use of object detection, speech to text and ESP32-CAM

TecNM: Tecnológico de Estudios Superiores de Jocotitlán, José Enrique Domínguez-Nava, Adriana Reyes-Nava, TecNM: Tecnológico de Estudios Superiores de Jocotitlán, Leopoldo Gil-Antonio, TecNM: Tecnológico de Estudios Superiores de Jocotitlán · ECORFAN eBooks · 2024

This work presents the design of an assistant for people with visual disabilities, which combines realtime object detection and the use of voice recognition to facilitate verbal interactions. The project is divided into two phases: the first develops a server in Python that receives user requests using a keyword. Upon detection, the server sends an HTTP request to the ESP32-CAM, which captures images of the environment. These are processed with YOLOv8 to identify the objects present. The server then generates a response in audio format, played through Bluetooth headphones, describing the detected objects. The second phase seeks to improve object detection and help the user find a particular one, indicating its location or distance. The work focuses on the first phase, which covers the design of the communication between the user, the server and the ESP32-CAM.

Read the paper · More papers on PaperTik