Using sub-word n-gram models for dealing with OOV in large vocabulary speech recognition for Latvian

Askars Salimbajevs, Jevgenijs Strigins · DSpace repository (University of Tartu) · 2015

In the Latvian language, one word can have tens or even hundreds of surface forms.This is a serious problem for large vocabulary speech recognition.Inclusion of every form in vocabulary will make it intractable, but, on the other hand, even with a vocabulary of 400K, the out-ofvocabulary (OOV) rate will be very high.In this paper, the authors investigate the possibility of using sub-word vocabularies where words are split into frequent and common parts.The results of our experiment show that this allows to significantly reduce the OOV rate.

Read the paper · More papers on PaperTik