Visual text-to-auditory conversion system

Archana Yashodip Chaudhari, Devanshu Kamdi, Romila Katkar, Nishant Pawar, Aradhay Kopulwar · 2025

The project focuses on providing an integrated image-to-text-to-speech system capable of embedding OCR and TTS technology in a user-friendly interface. The process involves taking an image through a standard camera, processing it by a grayscale conversion, and then text extraction using Tesseract OCR. Once the text is extracted, the eSpeak engine will be applied to produce synthetic speech to read it aloud. An easy-to-use graphical user interface (GUI) based on Tkinter would allow capturing the image with a click of a button and providing feedback in real-time about the success or failure of the operation. Errors will be handled to ensure that information on errors is provided to the user for the running of the system. This has been developed essentially for the benefit of the visually impaired or people with reading disabilities to provide an efficient automated solution for converting written text into spoken words.

Read the paper · More papers on PaperTik