Training time reduction and performance improvements from multilingual techniques on the BABEL ASR task

Sebastian Stüker, Markus Müller, Quốc Bảo Nguyễn, Alex Waibel · 2014

In the IARPA sponsored program BABEL we are faced with the challenge of training automatic speech recognition systems in sparse data conditions in very little time. In this paper we show that by using multilingual bootstrapping techniques in combination with multilingual deep belief bottle neck features that are only fine tuned on the target language the training time of an LVCSR system can be essentially halved while the word error rate stays the same. We show this for recognition systems on Tagalog, making use of multilingual systems trained on the other four languages of the Babel base period: Cantonese, Pashto, Turkish, and Vietnamese.

Read the paper · More papers on PaperTik