IndicST: Indian Multilingual Translation Corpus For Evaluating Speech Large Language Models

Sanket B. Shah, Kavya Ranjan Saxena, Kancharana Manideep Bharadwaj, Sharath Adavanne, Nagaraj Adiga · 2025

The integration of speech modalities into large language models, known as Speech LLMs, is a promising area of research for applications like automatic speech recognition (ASR) and automatic speech translation (AST). While several datasets exist for ASR, there is a critical gap for AST tasks in Indian languages. To fill this gap, we introduce IndicST, a new dataset tailored for training and evaluating Speech LLMs for AST tasks (including ASR and TTS), featuring meticulously curated, automatically and manually verified synthetic data. The dataset offers 10.8k hrs of training data and 1.13k hrs of evaluation data. Additionally, we present a scalable data collection methodology that allows easy development for other speech tasks. Our analysis examines the performance of various Speech LLMs on ASR and AST tasks, accompanied by insightful findings. We also explore the performance of these models across different prompts, highlighting the significant potential of our research in enhancing ASR and AST capabilities for Indian languages.

Read the paper · More papers on PaperTik