Spoken language identification using score vector modeling and support vector machine
Ming Li, Hongbin Suo, Xiao Guang Wu, Ping C. Lu, Yonghong Yan · 2007
The support vector machine (SVM) framework based on generalized linear discriminate sequence (GLDS) kernel has been shown effective and widely used in language identifica-tion tasks. In this paper, in order to compensate the distortions due to inter-speaker variability within the same language and solve the practical limitation of computer memory requested by large database training, multiple speaker group based discrim-inative classifiers are employed to map the cepstral features of speech utterances into discriminative language characterization score vectors (DLCSV). Furthermore, backend SVM classifiers are used to model the probability distribution of each target language in the DLCSV space and the output scores of back-end classifiers are calibrated as the final language recognition scores by a pair-wise posterior probability estimation algorithm. The proposed SVM framework is evaluated on 2003 NIST Lan-guage Recognition Evaluation databases, achieving an equal er-ror rate of 4.0 % in 30-second tasks, which outperformed the state-of-art SVM system by more than 30 % relative error re-duction. Index Terms: spoken language identification, support vector machine, score vector modeling