Experimental improvements of a language Id system

Kung-Pu Li · 2002

Previously, automatic language identification systems provided good results by using syllabic "on-set" spectral features; they identified languages by finding the "nearest match" speakers who were closet to the test utterance. The present authors we show that augmenting the training data by adding speakers achieves a better gender balance in the data and reduces the error rate by more than 10%. Adding features like syllabic "coda" and "prosodic" features show very different results which can then be merged with the syllabic "on-set" spectral features to reduce errors an additional 10%. A dimensionality reduction by means of the principal components shows not only a reduction in computation and memory requirements, but also improves language identification performance when the eigenvectors are normalized with different weights. The combination of all these factors yields a significant improvement in performance when compared with the previous baseline system.

Read the paper · More papers on PaperTik