Statistical Language Modeling for Automatic Speech Recognition of Agglutinative Languages

Ebru Arsoy, Mikko Kurimo, Murat Saralar, Teemu Hirsimki, Janne Pylkknen, Tanel Alume, Haim Sak · InTech eBooks · 2008

Speech Recognition, Technologies and Applications 194the OOV rate is less than 1% for North American Business news (Rosenfeld, 1995).Highly inflectional and agglutinative languages suffer from high number of OOV words with similar size vocabularies.In our Turkish BN transcription system, the OOV rate is 9.3% for a 50K lexicon.For other agglutinative languages like Finnish and Estonian, OOV rates are around 15% for a 69K lexicon (Hirsimäki et al., 2006) and 10% for a 60K lexicon respectively and 8.27% for Czech, a highly inflectional language, with a 60K lexicon (Podvesky & Machek, 2005).As a rule of thumb an OOV word brings up on average 1.5 recognition errors (Hetherington, 1995).Therefore solving the OOV problem is crucial for obtaining better accuracies in the ASR of agglutinative languages.OOV rate can be decreased to an extent by increasing the vocabulary size.However, even doubling the vocabulary is not a sufficient solution, because a vocabulary twice as large (120K) would only reduce the OOV rate to 6% in Estonian and 4.6% in Turkish.In Finnish even a 500K vocabulary of the most common words still gives 5.4% OOV in the language model training material.In addition, huge lexicon sizes may result in confusion of acoustically similar words and require a huge a m o u n t o f t e x t d a t a f o r r o b u s t l a n g u a g e model estimates.Therefore, sub-words are proposed as language modeling units to alleviate the OOV and data sparseness problems that plague systems based on word-based recognition units in agglutinative languages.In sub-word-based ASR; (i) words are decomposed into meaningful units in terms of speech recognition, (ii) these units are used as vocabulary items in n-gram language models, (iii) decoding is performed with these n-gram models and sub-word sequences are obtained, (iv) word-like units are generated from sub-word sequences as the final ASR output.In this chapter, we mainly focus on the decomposition of words into sub-words for LVCSR of agglutinative languages.Due to inflections, ambiguity and other phenomena, it is not trivial to automatically split the words into meaningful parts.Therefore, this splitting can be performed by using rule-based morphological analyzers or by some statistical techniques.The sub-words learned with morphological analyzers and statistical techniques are called grammatical and statistical sub-words respectively.Morphemes and stem-endings can be used as the grammatical sub-words.The statistical sub-word approach presented in this chapter relies on a data-driven algorithm called Morfessor Baseline (Creutz & Lagus, 2002;Creutz & Lagus, 2005) which is a language independent unsupervised machine learning method to find morpheme-like units (called statistical morphs) from a large text corpus.After generating the sub-word units, n-gram models are trained with sub-words similarly as if the language modeling units were words.In order to facilitate converting sub-word sequences into word sequences after decoding, word break symbols can be added as additional units or special markers can be attached to non-initial sub-words in language modeling.ASR systems that successfully utilize the n-gram language models trained for sub-word units are used in the decoding task.Finally, word-like ASR output is obtained from sub-word sequences by concatenating the sub-words between consecutive word breaks or by gluing marked non-initial sub-words to initial ones.The performance of words and sub-words are evaluated for three agglutinative languages, Finnish, Estonian and Turkish.This chapter is organized as follow: In Section 2, our statistical language modeling approaches are explained in detail.Section 3 contains the experimental setup for each language.Experimental results are given in Section 4. Finally, this chapter is concluded with a detailed comparison of the proposed approaches for agglutinative languages.www.intechopen.

Read the paper · More papers on PaperTik