Spoken Language Identification Based on I-vectors and Conditional Random Fields
Panikos Heracleous, Yasser F. O. Mohammad, Koichi Takai, Keiji Yasuda, Akio Yoneyama · 2018
The task of an automatic language identification (LID) system is to automatically identify the language in a spoken utterance. Language identification can be applied as front-end to speech-to-speech translation systems, in speaker diarization, and at call centers to automatically route incoming calls to appropriate native speaker operators. In the current study, a method for automatic language identification based on i-vector paradigm and conditional random fields (CRF) is presented. CRF belong to discriminative classifiers and use an exponential distribution to model a sequence given the observation sequence. This allows non-independent observations, and allows also non-local dependencies between state and observation. When the proposed method is evaluated on the NIST 2015 i-vector Machine Learning Challenge task for the recognition of 50 in-set languages, a 3.7% equal error rate (EER) (i.e., miss probability equal to false alarms) is achieved. Using support vector machines (SVM) a 5.2% EER, using probabilistic linear discriminant analysis (PLDA) a 6.7% EER, and when using convolutional neural networks (CNN) a 4.2% EER are achieved.