Multilingual Automatic Speech Recognition for Indian Languages-E2E Framework
R. Geetha Rajakumari, D. Karthika Renuka, L. Ashok Kumar, C Thiraviya, S Vaimitra, Sathishkumar Veerappampalayam Easwaramoorthy · 2024
Automatic Speech Recognition (ASR) technology has made significant strides in supporting a wide array of languages, with current estimates suggesting coverage for approximately 100 to 150 languages globally. These languages include major linguistic entities such as English, Spanish, Chinese (Mandarin), French, German, Japanese, and Arabic, among others. The focus of this research work is to develop a robust multilingual ASR system tailored specifically for Tamil and English languages, acknowledging the linguistic diversity and code-switching tendencies prevalent in Indian speech patterns. One of the primary challenges addressed by this research work is the accurate identification and segmentation of languages within mixed-language audio. The proposed work is done by creating an end- to-end (E2E) framework using a Seq2Seq model, a neural network architecture well-suited for sequence generation tasks such as speech transcription. The system achieves an impressive accuracy rate of 94%.