Automatic Language Identification using Machine learning Techniques

Hariraj Venkatesan, T. Varun Venkatasubramanian, J. Sangeetha · 2018

Investigation in the area of spoken language identification on regional languages aids to broaden the outreach of technology to regional language speakers and also gives to the preservation of regional languages. In this paper, we report our work on identifying spoken data in four local Indian languages Kannada, Hindi, Tamil and Telugu. Automatic Language Identification systems take a speech signal as input and perform computations on the speech input to classify it into one of the natural languages. Mathematical computations performed on the properties of a speech signal such as frequency or amplitude can be used to derive information about the audio and its speaker. In this paper, Mel Frequency Cepstral Coefficients (MFCC) has been used to derive features of speech signals that can be used for identifying languages. For classification purposes, Support Vector Machines and Decision Tree classifiers were used and we got accuracies of 76% and 73% respectively.

Read the paper · More papers on PaperTik