Investigation of Handwritten Image-To-Speech Using Deep Learning
Manju S, J Anitha, Sujitha Juliet D · 2024
In the realm of accessibility technology, this paper introduces a pioneering method for converting handwritten images to speech. The work primarily focuses on recognizing handwritten text and subsequently converting it into audible speech. The work comprises two fundamental phases: initially, a deep learning algorithm is employed to recognize and extract text from handwritten images precisely. Following this, the recognized text is converted into speech using the pyttsx3 library which offers flexibility in voice selection, speech rate, and other parameters, making it highly customizable for different applications. To ascertain the most effective algorithm for the text recognition phase, a comparative analysis is conducted among the top three algorithms-Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and Transformer-based models - based on their accuracy in recognizing handwritten text. Experimental results indicate that CNNs surpass the other two algorithms, achieving an accuracy of 92%. This work offers a pragmatic approach to converting handwritten content into speech, with potential applications in aiding the visually impaired and enhancing accessibility to written information.