A Chinese-Japanese Parallel Corpus for Neural Machine Translation Based on Web-Crawled News Data

Zhonghui Gao, Jinyi Zhang, Ye Tian, Tadahiro Matsumoto · 2024

Current research on Japanese-Chinese parallel corpora is limited, particularly in the general field, and existing corpora are often small in scale. The purpose of this study is to fill this gap by constructing a large scale of Chinese-Japanese parallel corpus, sourced from web-crawled online news data. We also evaluated the quality of the constructed Web Crawled Corpus of Japanese and Chinese News (WCC-JCN), which comprised approximately 660K Chinese-Japanese sentence pairs, by calculating its BLEU scores and comparing it with the established ASPEC-JC corpus. The WCC-JCN is made freely available for download, provided it is used for research purposes only, thereby contributing to future research in this area.

Read the paper · More papers on PaperTik