Speech Recognition and Transcription

Tanushree Borase, R Thamizhamuthu · 2025

In this paper, we propose a novel speech-to-text model. This is a user-friendly platform designed to convert spoken text into written text efficiently and accurately. It allows users to record their speech, transcribe in real time and saves output in a downloadable format such as pdf. It offers features such as language selection and real-time display of transcribed text. By supporting multiple languages such as (English, Hindi, Marathi, Tamil, Bengali), it addresses the needs of a global audience. Overall, this offers a versatile and accessible solution for converting speech to text, making it a valuable tool for users across different fields. This paper details the system’s architecture, methodologies employed, and the results of extensive testing across different languages and audio qualities. The system’s impact on accessibility and productivity in multilingual environments is discussed, with potential areas for future enhancement outlined.

Read the paper · More papers on PaperTik