IMPROVED TEXT NORMALIZATION AND LANGUAGE MODELS FOR SPEED'S AUTOMATIC SPEECH RECOGNITION SYSTEM

Cristian Manolache, Alexandru-Lucian Georgescu, Horia Cucu, Verginica Barbu Mititelu, Corneliu Burileanu · 2020

Automatic speech recognition (ASR) systems that use word-based language models require periodical updates to include new named entities (e.g. coronavirus, COVID-19) or collocations. Moreover, in particular for the Romanian language, the new hyphenated words pose additional problems. In this context, our study presents SpeeD's efforts in collecting new text corpora and using them for language modelling in the context of ASR. We also present the improvements made in the text normalization module to address the problems posed by hyphenated words. We evaluate the resulting language models both in terms of their ability to predict future words (perplexity and out-of-vocabulary rate) and in terms of their usefulness in ASR (word error rate). We report ASR relative improvements of around 10% for spontaneous speech, with small degradations for read speech.

Read the paper · More papers on PaperTik