Study of different backends in a state-of-the-art language recognition system

Mikel Peñagarikano, Amparo Varona, Mireia Díez, Luis Javier Rodríguez-Fuentes, Germán Bordel · 2012

State of the art language recognition systems usually add a backend prior to the linear fusion of the subsystems scores. The backend plays a dual role. When the set of languages for which models have been trained does not match the set of target lan-guages, the backend maps the available scores to the space of target languages. On the other hand, the backend serves as a precalibration stage that adapts the space of scores. In this work, well known backends (Generative Gaussian Backend, Discrim-inative Gaussian Backend and Logistic Regression Backend) and newer proposals (Fully Bayesian Gaussian Backend and Gaussian Mixture Backend) are analyzed and compared. The effect of applying a T-Norm or a ZT-Norm is also analyzed. Fi-nally the effect of discarding development signals, those with the highest scores, is also studied. Experiments have been car-ried out on the NIST 2009 LRE database, using a state-of-the-art Language Recognition System consisting of the fusion of five subsystems: A Linearized Eigenchannel GMM (LE-GMM) subsystem, an iVector subsystem and three phone-lattice-SVM subsystems. Best performance was attained by Gaussian Mix-ture Backend (1.25 EER), yielding 23 % relative improvement with respect to the baseline (1.62 EER).

Read the paper · More papers on PaperTik