Sound Sight-A TTS and STT Integrated Model Using OCR and ASR

Deepak Kumar A, A. S., R Aravindan, Anumeera Balamurali · 2024

This paper discusses the feasibility of using an automated pipeline that is text-to-speech and speech-to-text, which will be powered with the usage of Optical Character Recognition (OCR) that may be used for extracting text from images or documents and Automatic Speech Recognition (ASR) which can be used for the benefit of using algorithms and machine learning techniques to give more accurate results. The paper also explores the system architecture, with also a look at the methodology as it implements the technologies.

Read the paper · More papers on PaperTik