Dialect Identification of the Bengali Language
Elizabeth Behrman, Arijit Santra, Siladitya Sarkar, Prantik Roy, Ritika Yadav, Soumi Dutta, Arijit Ghosal · 2021
Dialect is a distinguishable form or regional variation of a single language specific to a geographic region or social group. Dialects of a language are the variation in grammar and vocabulary of the same language, because of the geographical and ethnic differences of the speakers. All languages have a large number of dialects within it. Dialects also indicate cultural variation specific to a geographic region. Dialect identification is becoming a significant and promising research area as it is finding its importance in improvement of performance in several modern application domains like voice-controlled smart devices, security domain, speech-to-text conversion, e-health, and e-learning. Bengali is such a language which is spoken not only in India but also spoken by a large number of people outside India. Bengali language is the second most widely spoken language in India after Hindi. Bengali is primarily spoken in West Bengal, Tripura, and Assam’s Barak Valley in India along with several other geographic locations throughout India also. Because of this large geographic span, there exist several dialects in Bengali language. India is the land of versatility. Several languages are being spoken in India and each of these languages has many dialects. As dialect identification is a promising research area; several research works are already carried to identify dialects of Indian languages until now. Most of these works involved dialect identification of Hindi language as Hindi is the most widely spoken language in India. Comparatively less works have been carried out to identify dialects of Bengali language. There exist mainly six dialects in Bengali language—Rarhi, Bangali, Manbhumi (Jharkhandi), Varendri, Rangpuri, and Sundarbani. This work considers three dialects—Rarhi, Bangali, and Manbhumi (Jharkhandi). These dialects are prominent dialects of Bengali language spoken in western, southern, and eastern regions of West Bengal. Any dialect can be identified based on phonemes and pronunciation of a speaker. Also loudness, tonality, and nasality play important roles while identifying dialects of a certain language. All these characteristics can be best measured through time domain and frequency domain aural features. Rarhi, Bangali, and Manbhumi (Jharkhandi) dialects also vary in time domain and frequency domain. Zero Crossing Rate (ZCR) is a time domain aural feature which defines the number of times the aural signal changes its sign that is the number of times it crosses the zero axis. Due to variation of loudness and nasality, the ZCR value of three dialects differs. Mel Frequency Cepstral Coefficients (MFCCs) are frequency domain aural feature which is mostly used for speech processing and speaker recognition which involves capturing the characteristics of phonemes, tonality, and pronunciation of a speaker. Skewness, a frequency domain aural feature, is applied in this exertion to incarcerate the certain breed of unlikeness among the three dialects. Hence this advised approach banks on the said time domain and frequency domain aural features. Neural Network (NN), Naïve Bayes, Random Forest, and K-Nearest Neighbor classifiers have been exercised for identification tasks. Experimental results have been compared with existing methodologies to reflect the efficiency of the proposed system.