Spoken Languages Identification for Indian Languages in Real World Condition
Sujeet Kumar, H Muralikrishna, Veena Thenkanidiyoor, A. D. Dileep · 2024
This work uses deep learning and advanced audio features to detect Indian spoken languages. Using pre-trained models such as wav2vec, data2vec, and ccc-wav2vec, we retrieved the feature representations of audio. Spoken language identification models were trained independently on each feature representation. To achieve this, an utterance-level embedding called u-vector with WSSL (within-sample similarity loss) is trained along with a simple DNN (Deep Neural Network) classifier on these features. In this paper, 12 Indian-spoken languages (including English) are considered and trained for only 10 hours of speech data from each language. The results show that using these feature representations and utterance-level embedding, a simple DNN can efficiently identify different Indian languages.