Recent advances in ASR applied to an Arabic transcription system for Al-Jazeera
Patrick Cardinal, Ahmed M. Abdelatty Ali, Najim Dehak, Yu Zhang, Tuka Alhanai, Yifan Zhang, James Glass, Stephan Vogel · 2014
This paper describes a detailed comparison of several state-of-the-art speech recognition techniques applied to a limited Ara-bic broadcast news dataset. The different approaches were all trained on 50 hours of transcribed audio from the Al-Jazeera news channel. The best results were obtained using i-vector-based speaker adaptation in a training scenario using the Min-imum Phone Error (MPE) criteria combined with sequential Deep Neural Network (DNN) training. We report results for two different types of test data: broadcast news reports, with a best word error rate (WER) of 17.86%, and a broadcast conver-sations with a best WER of 29.85%. The overall WER on this test set is 25.6%. Index Terms: Arabic, ASR system, Kaldi 1.