Experiments with LVCSR based language identification

Tanja Schultz, Ivica Rogina, Alexander H. Waibel · Repository KITopen (Karlsruhe Institute of Technology) · 1995

Automatic language identification is an important problem in building multilingual speech recognition and understanding systems. We have developed a front-end LID module based on LVCSR to identify English, German, and Spanish language for use in spontaneous speech-to-speech translation. We studied the constitution of different levels of knowledge to identify a language, i.e. the phonetic, phonotactic, lexical, and syntactic-semantic knowledge. A comparison of LID systems using different levels of these knowledge sources is presented. We showed that the incorporation of lexical and linguistic knowledge leads to a reduction of the language identification error by up to 50%. 1. INTRODUCTION In recent years language identification (LID) has received renewed and increased interest as LVCSR technology is being applied to multiple languages. The arrival of multilingual databases like the OGI corpus [1], [2] and the Spontaneous Scheduling Task (SST) [3] enable us to compare different approac...

Read the paper · More papers on PaperTik