Constructing parallel corpora for six indian languages via crowdsourcing

Matt Post, Chris Callison-Burch, Miles Osborne · 2012

Recent work has established the efficacy of Amazon’s Mechanical Turk for constructing parallel corpora for machine translation re-search. We apply this to building a collec-tion of parallel corpora between English and six languages from the Indian subcontinent: Bengali, Hindi, Malayalam, Tamil, Telugu, and Urdu. These languages are low-resource, under-studied, and exhibit linguistic phenom-ena that are difficult for machine translation. We conduct a variety of baseline experiments and analysis, and release the data to the com-munity. 1

Read the paper · More papers on PaperTik