Incremental and Flexible Extraction of Parallel Corpus from the Web
Xiaofeng Liu, Yucheng Zhen, Dongyang Li · 2024
Extracting parallel corpus from the web at scale is important for machine translation and other multilingual processing tasks. The paper proposes an incremental and flexible web parallel corpus extraction approach, which incrementally updates language text length statistics for domains by continuously downloading, scanning and analyzing Common Crawl's web crawling archives, and extracting parallel corpus. For any given interested language pairs, the web sites to be crawled are determined based on language text length statistics for domains and crawled according to the target language pairs, and non-target domains and links are discarded. The paper also proposes a new intermediate method for sentences alignment, which globally aligns sentences based on semantic similarity within a multilingual domain. Experiments have shown that: (1) our extraction method can continuously and flexibly extract interested parallel corpus via specifying target language pairs; (2) the proposed intermediate method is significantly better than the global method in terms of alignment efficiency, and can complete some alignments that cannot be done by the local method; (3) out of 6 language directions, the extracted parallel corpora are superior to existing web open source parallel corpora in 4 low-medium resource directions and close to the best available web open source parallel corpus in 2 high-resource directions.