The AMARA Corpus: Building Resources for Translating the Web's Educational Content
Francisco Guzmán, Hassan Sajjad, Vogel, Stephan, Ahmed Abdelalí · 2013
In this paper, we introduce a new parallel corpus of subtitles of educational videos: the AMARA corpus for online edu-cational content. We crawl a multilingual collection com-munity generated subtitles, and present the results of pro-cessing the Arabic–English portion of the data, which yields a parallel corpus of about 2.6M Arabic and 3.9M English words. We explore different approaches to align the seg-ments, and extrinsically evaluate the resulting parallel corpus on the standard TED-talks tst-2010. We observe that the data can be successfully used for this task, and also observe an absolute improvement of 1.6 BLEU when it is used in com-bination with TED data. Finally, we analyze some of the specific challenges when translating the educational content. 1.