Integration of Computer Vision for Book Detection and Text-to-Speech Conversion Using API for Blind People

Tirthak Likhar, Ameen Khan, Bharat Singh Parihar, Nitin Rakesh, Monali Gulhane, Pratik K. Agrawal · 2024

The visually impaired and blind community has devised various methods to access printed materials despite their visual challenges. However, current solutions such as braille, audiobooks, and human assistants suffer from limitations in availability, coverage, privacy, and independence. This paper introduces a book recognition system that integrates computer vision and machine learning technologies to serve individuals with low vision. By simply pointing the camera at books, the system can audio titles, summaries, and captions. The system focuses on uses OpenCV to extract text for the real-time book detection from Tesseract's optical character recognition (OCR) capabilities. Information extracted through text-to-speech conversion is presented as audio output, facilitating information retrieval. Employing user-centric and low-barrier approaches, this tool addresses the shortcomings associated with traditional methods. The system architecture thoughtfully incorporates next-generation technologies such as object recognition, OCR, and speech synthesis, seamlessly integrating them into a cohesive and userfriendly solution tailored for visually impaired individuals. This groundbreaking solution tackles the issue of accessing printed materials head-on, paving the way for independent living and equal opportunities. An extensive literature review explores foundational concepts in computer vision, object recognition, shape matching, and multimodal classification.

Read the paper · More papers on PaperTik