Harvesting Parallel Text in Multiple Languages with Limited Supervision

Luciano Barbosa, Vivek Kumar Rangarajan Sridhar, Mahsa Yarmohammadi, Srinivas Bangalore · International Conference on Computational Linguistics · 2012

The Web is an ever increasing, dynamically changing, multilingual repository of text. There have been several approaches to harvest this repository for bootstrapping, supplementing and adapting data needed for training models in speech and language applications. In this paper, we present semi-supervised and unsupervised approaches to harvesting multilingual text that rely on a key observation of link collocation. We demonstrate the eectiveness of our approach in the context of statistical machine translation by harvesting parallel texts and training translation models in 20 dierent languages. Furthermore, by exploiting the DOM trees of parallel webpages, we extend our harvesting technique to create parallel data for resource limited languages in an unsupervised manner. We also present some interesting observations concerning the socio-economic factors that the multilingual Web reflects.

Read the paper · More papers on PaperTik