Automatic Acquisition of Parallel Corpora from Websites with Dynamic Content
Yulia Tsvetkov, Shuly Wintner · 2010
Parallel corpora are indispensable resources for a variety of multilingual natural language processing tasks.This paper presents a technique for fully automatic construction of constantly growing parallel corpora.We propose a simple and effective dictionary-based algorithm to extract parallel document pairs from a large collection of articles retrieved from the Internet, potentially containing manually translated texts.This algorithm was implemented and tested on Hebrew-English parallel texts.With properly selected thresholds, precision of 100% can be obtained.