Phonetic-context mapping in language identification

Jiří Navrátil, Werner Zühlke · 1997

This paper deals with the problem of exploiting information from a wide phonetic context for the purpose of language identification. Two approaches to language modeling are presented here: 1) modified bigrams with a context -mapping matrix and 2) language models based on binary decision trees. Both models were incorporated in a phonotactic language identifier with a double-bigram decoding architecture and were shown to consistently improve the performance of standard bigrams. Measured on the NIST'95 evaluation set, the described system outperforms the state-of-the-art phonotactic components and is, at the same time, computationally less expensive. 1. INTRODUCTION Automatic language identification (ALI) is a task of recognizing the language from a spoken test sentence. The ability of machines to distinguish between different languages becomes important with the trend in globalizing communication technology and providing wide multilingual services. Besides other solutions for ALI based ...

Read the paper · More papers on PaperTik