Diacritization, automatic segmentation and labeling for Levantine Arabic speech
Yousef Ajami Alotaibi, Ali Hamid Meftah, Sid‐Ahmed Selouani · 2013
It is generally acknowledged that a reliable speech corpus is necessary for any application involving speech processing. In this paper, we propose methods to improve the BBN/AUB DARPA Babylon Levantine Arabic speech corpus to increase its reliability and efficiency. For this purpose, correction of pronunciation, diacritization, and new transcription are performed manually along with automatic phoneme segmentation and labeling. The comparison with the original transcription of the corpus shows a clear improvement in the output results.