Automatic language identification for seven Indian languages using higher level features
Chithra Madhu, Anu George, Leena Mary · 2017
This paper proposes an approach for automatic language identification (LID) for seven Indian languages. The proposed system uses language dependent phonotactic features and prosodic information. Phonetic Engine (PE) which serves as the front end of the phonotactic based LID system converts input speech utterance to a sequence of phonetic symbols. Syllable boundaries are detected and phones within a syllable boundary are grouped and phono-tactic rules are applied to get syllables. Two consecutive syllables are numerically represented to get phonotactic feature vectors. Prosodic feature vectors are obtained by concatenating features of three consecutive syllables. A multilayer feed forward neural network (NN) classifier is used at the back-end for language identification. The ANN classifier is trained with two hour duration data from each of the seven languages. Target languages include Bengali, Hindi, Telugu, Urdu, Assamese, Punjabi and Manipuri.