Robust & Compact End-to-End Hindi Language ASR System
Hruturaj Nikam, Mahesh Bhargava, Pavan Dhote, Sanhita Patil, Sibadatta Sasmal, Lenali Singh, Swati Mehta, Ajai Kumar · 2022
Deep neural network architectures are highly in-corporated in modern (progressive) automatic speech recognition(ASR) systems which outperforms conventional and hybrid systems. Even after adapting to such system we still facing a problem to make balance between accuracy and model size (no. of trainable parameters). We present our experiment in building a robust encoder-decoder based end-to-end ASR model for the Hindi language. Starting with, we have trained the QuartzNet model with two different architectures QuartzNet-5x5 and QuartzNet-15x5 by keeping the same training parameters and dataset. The speech corpus consists of 2648.9 hours of Hindi labelled audio data collected from various domains and sources. We have analyzed the performance of both models and observed that QuartzNet-15x5 has a radical improvement of 24% in accuracy. For building and training the models, we have used the PARAM SIDDHI AI system and training recipes from the open source NeMo toolkit.