University of the Basque Country (EHU) Systems for the 2011 NIST Language Recognition Evaluation

Mikel Peñagarikano, Amparo Varona, Luis Javier Rodríguez-Fuentes, Mireia Díez, Germán Bordel · 2011

This paper describes the systems developed by the Software Technologies Working Group (http://gtts.ehu.es) of the University of the Basque Country for the 2011 NIST Language Recognition Evaluation. Four different systems (one primary and three contrastive) were submitted, consisting of a fusion of five subsystems: a Linearized Eigenchannel GMM (LE-GMM) subsystem, an iVector subsystem and three phone-lattice-SVM subsystems based on the publicly available BUT decoders for Czech, Hungarian an Russian. The four submitted systems were identical except for the backend approach and the development dataset used to estimate the backend and fusion parameters. Multiclass fusion was performed separately for each nominal duration. A development set was defined, including the evaluation sets of LRE07 and LRE09 and the development data provided by NIST for 9 additional languages in LRE11. Systems were evaluated on 10 random partitions of the development set, using one half for estimating backend and fusion parameters and the other half for testing. The average cost as defined in the LRE11 evaluation plan was used as performance measure. The primary system yielded an actual average cost of 0.038 (±0.002), being Hindi-Urdu, by far, the most challenging pair, with an actual average cost of 0.222.

Read the paper · More papers on PaperTik