Identification of top-3 spoken Indian languages: An Ensemble learning-based approach
Himadri Mukherjee, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy · 2018
Speech recognition has developed considerably for English but there has not been much development in Indic languages. Speech Recognition in Indic languages is itself challenging which complicates even more in multilingual scenario. There is a pressing need for Indic speech recognition systems and a fully functional variant of the same is yet to be developed. One reason for this is the multi lingual nature of our country in addition to the complexity of the Indic languages. It is very much important to identify the language specific segments from multi lingual speech before attempting recognition. In this paper, we have presented a system to segregate the top 3 spoken languages in India encompassing English, Hindi and Bangla. We have experimented with segregation of Bangla alone from the 3 languages as well driven by the motivation that Bangla is our mother tongue. Experiments were performed on more than 24 hours of data and highest accuracies of 97.13% and 96.44% has been obtained in segregating Bangla from the rest and trilingual segregation respectively with MFCC-based features coupled with Ensemble learning-based classification.