Using cross-decoder co-occurrences of phone n-grams in SVM-based phonotactic language recognition
Mikel Peñagarikano, Amparo Varona, Luis Javier Rodríguez-Fuentes, Germán Bordel · 2010
Most common approaches to phonotactic language recognition deal with several independent phone decoders. Decodings are processed and scored in a fully uncoupled way, their time align-ment (and the information that may be extracted from it) being completely lost. Recently, we have presented a new approach to phonotactic language recognition which takes into account time alignment information, by considering cross-decoder co-occurrences of phones or phone n-grams at the frame level. Ex-periments on the NIST LRE2007 database demonstrated that using co-occurrence statistics could improve the performance of baseline phonotactic recognizers. In this work, the approach based on cross-decoder co-occurrences of phone n-grams is fur-ther developed and evaluated. Systems were built by means of open software (Brno University of Technology phone de-coders, LIBLINEAR and FoCal) and experiments were carried out on the NIST LRE2007 database. A system based on co-occurrences of phone n-grams (up to 4-grams) outperformed the baseline phonotactic system, yielding around 8 % relative improvement in terms of EER. The best fused system attained 1,90 % EER (a 16 % improvement with regard to the baseline system), which supports the use of cross-decoder dependencies for improved language modeling.