Coping with out-of-vocabulary words: Open versus huge vocabulary asr

Matteo Gerosa, Marcello Federico · 2009

This paper investigates methods for coping with out-of-vocabulary words in a large vocabulary speech recognition task, namely the automatic transcription of Italian broadcast news. Two alternative ways for augmenting a 64 K(thousand)-word recognition vocabulary and language model are compared: introducing extra words with their phonetic transcription up to 1.2 M (million) words, or extending the language model with so-called graphones, i.e. subword units made of phone-character sequences. Graphones and phonetic transcriptions of words are automatically generated by adapting an off-the-shelf statistical machine translation toolkit. We found that the word-based and graphone based extentions allow both for better recognition performance, with the former performing significantly better than the latter. In addition, the word-based extension approach shows interesting potential even under conditions of little supervision. In fact, by training the grapheme to phoneme translation system with only 2 K manually verified transcriptions, the final word error rate increases by just 3% relative, with respect to starting from a lexicon of 64 K words.

Read the paper · More papers on PaperTik