Crowdsourcing High-Quality Parallel Data Extraction from Twitter

Wang Ling, Luís Marujo, Chris Dyer, Alan W. Black, Isabel M. Trancoso · 2014

High-quality parallel data is crucial for a range of multilingual applications, from tuning and evaluating machine translation systems to cross-lingual annotation projection.Unfortunately, automatically obtained parallel data (which is available in relative abundance) tends to be quite noisy.To obtain high-quality parallel data, we introduce a crowdsourcing paradigm in which workers with only basic bilingual proficiency identify translations from an automatically extracted corpus of parallel microblog messages.For less than $350, we obtained over 5000 parallel segments in five language pairs.Evaluated against expert annotations, the quality of the crowdsourced corpus is significantly better than existing automatic methods: it obtains an performance comparable to expert annotations when used in MERT tuning of a microblog MT system; and training a parallel sentence classifier with it leads also to improved results.The crowdsourced corpora will be made available in http://www.cs.cmu.edu/~lingwang/microtopia/.

Read the paper · More papers on PaperTik