Language Detection Based on Audio for Indian Languages
A. M. Amogh, Anshu Priya, Thanvitha Sai Kanchumarti, Likhitha Ram Bommilla, Rajeshkannan Regunathan · 2024
The Indian subcontinent has a varied linguistic community, with 22 officially recognized languages including countless dialects. Each language has its own distinct accent and dialect, making it difficult to determine the language spoken in a given nation. As a result, in such instances, the task of the spoken language identification (SLID) is extremely difficult. The main objective of this chapter is to tackle the problem by presenting a deep learning model that can correctly identify different Indian languages while expanding the number of languages that can be identified. This paper suggests a model for recognizing various Indian languages such as Hindi, Kannada, Bengali, Gujarati, Tamil, Telugu, Marathi, Malayalam, Punjabi, and Urdu. These languages were chosen because they are extensively spoken in India, and a few of them are similar to one another, and the proposed model can predict those similar languages correctly. The audio files are fed into the model in this chapter, which then preprocesses them to produce a spectrogram graph of the speech signals. Spectrogram graphs are important for representing audio signals because they provide information about the signal's time-varying frequency content. Following that, the mel-frequency cepstral coefficients (MFCC) technique is well used to extract the relevant features needed for identification of the language. In speech processing, MFCC is a common feature extraction method that converts an audio signal into a series of feature vectors. The feature vectors record the most important aspects of the speech signal, such as its frequency content and dynamics. After extracting the MFCC features, the deep learning model, particularly a convolutional neural network (CNN), is used to classify the input into various languages. The provided dataset is utilized to train the proposed model, which is subsequently assessed for its accuracy. The CNN model was selected because it has been demonstrated to be successful in speech and language processing applications such as language identification. The approach suggested in this chapter has several advantages. For starters, it increases the number of languages that can be correctly recognized, which is critical for speech and language processing applications in India. Second, using a deep learning model improves the job of language identification accuracy. Finally, the suggested model is scalable, which means it can be extended to other languages and dialects. Finally, this chapter suggests a deep learning-based approach to identifying spoken languages in India.