Automatic Recognition of Continuous Malayalam Speech using Pretrained Multilingual Transformers
Kavya Manohar, Gokul G. Menon, Ashish Abraham, Rajeev Rajan, A. R. Jayan · 2023
This article presents our work on the development of automatic speech recognition system for Malayalam language. Multilingual transformer models trained on self supervised learning technique has been fine-tuned for Malayalam language using publicly available annotated speech datasets. The datasets are diverse, which vary in the number of speakers, domain of speech and recording environments. We compare the results obtained by fine-tuning the transformer based models with traditional speech recognition approaches, in which the acoustic and language models are trained independently and then combined. The resulting word error rates vary widely with the test dataset under consideration. Our results point to the importance of considering pre-trained speech representations when developing ASR systems for languages with limited training resources. We make the model openly available on HuggingFace for reproducibility and adaptations to similar languages.