Indic-ST: A Large-Scale Multilingual Corpus for Low-Resource Speech-to-Text Translation
Nivedita Sethiya, Saanvi Nair, Puneet Walia, Chandresh Kumar Maurya · ACM Transactions on Asian and Low-Resource Language Information Processing · 2025
We introduce Indic-ST, a novel dataset for speech-to-text translation (ST) task from English to Indic languages to bridge the performance gap. ST involves converting spoken input in one language into written text in another, playing a key role in real-world applications like subtitling, lecture transcription, and multilingual communication systems. Despite several efforts like Meta’s seamless m4t, OpenAI’s Whisper, or Google USM model, the performance of ST models on low-resource languages lags to that of English (or high-resource languages like European languages). Indic-ST is compiled from four distinct domains: conversational audio, religious texts, education, and news, which combined results in the Indic-ST dataset. To the best of our knowledge, this is the largest low-resource ST data covering approximately 6,800 hours of English speech in the real human voice and text in 15 Indic languages with diverse scripts totaling approximately 900 GB in size. To assess the usefulness of the dataset, we present the baseline performance of individual language pairs using state-of-the-art ST models. We also present a unified multilingual English-to-Indic-ST model. The code and dataset are available at https://github.com/Nivedita5/Indic-ST .