Systematic Approach for Speech to Text Translation and Summarization
Satwinder Singh, P Aswin, Yash Paul, Shafiya Mushtaq, Rajesh Singh · Indian Journal of Science and Technology · 2025
Objectives: To develop a fully automated framework capable of accurately transcribing and translating native language speech into textual form without human intervention. Focusing on multilingual speech recognition, the research aims to evaluate system performance using real-world speech data, with particular emphasis on translation fidelity and contextual accuracy. Methods: The proposed framework processes raw audio input, exemplified by Indian Prime Minister Narendra Modi’s monthly "Mann Ki Baat (मन की बात)" addresses, through a multi-stage pipeline. Initially, the system converts speech into English text via automated speech recognition, followed by Hindi translation using the IBiLex extractor. Translation quality is rigorously assessed through similarity scoring across three dimensions: Topic Modeling, Context Modeling, and Combined Modeling, with performance benchmarks established through comparison against manually translated reference texts. Findings: Experimental results demonstrate that Context Modeling yields optimal performance, achieving consistent similarity scores of 0.0355 across all evaluation metrics (Precision@1, MRR, and Precision@all). When benchmarked against human translations, the framework attains a peak similarity score of 0.1905, revealing both the potential and current limitations of fully automated translation systems. These findings underscore the critical importance of contextual understanding in preserving semantic meaning during speech-to-text conversion. Novelty: This research contributes three significant advancements to the field of speech processing: First, it presents a novel end-to-end architecture for automated speech-to-text translation specifically designed for Indian language pairs. Second, it establishes "Mann Ki Baat (मन की बात)" as a valuable benchmark dataset for evaluating speech recognition in political and native-language contexts. Third, the study introduces a comprehensive evaluation methodology using similarity scoring across multiple modeling approaches, providing new insights into the relative importance of contextual versus topical information in machine translation systems. Keywords: Automatic Speech Recognition (ASR), Speech-to-text translation system (STT), Multi-lingual, Bilingual, Speech-to-text translation, Text summarization, Similarity score