Automatic Language Identification using Machine learning Techniques
Hariraj Venkatesan, T. Varun Venkatasubramanian, J. Sangeetha · 2018
Investigation in the area of spoken language identification on regional languages aids to broaden the outreach of technology to regional language speakers and also gives to the preservation of regional languages. In this paper, we report our work on identifying spoken data in four local Indian languages Kannada, Hindi, Tamil and Telugu. Automatic Language Identification systems take a speech signal as input and perform computations on the speech input to classify it into one of the natural languages. Mathematical computations performed on the properties of a speech signal such as frequency or amplitude can be used to derive information about the audio and its speaker. In this paper, Mel Frequency Cepstral Coefficients (MFCC) has been used to derive features of speech signals that can be used for identifying languages. For classification purposes, Support Vector Machines and Decision Tree classifiers were used and we got accuracies of 76% and 73% respectively.