A Chinese-Japanese Parallel Corpus for Neural Machine Translation Based on Web-Crawled Data from NetEase Cloud Music
Haowei Li, Jinyi Zhang, Ye Tian, Tadahiro Matsumoto · 2024
In machine translation in the field of natural language processing, film and TV subtitles are commonly used as corpora. However, their availability is limited, and the process of crawling and aligning these subtitles can be challenging. To overcome this limitation, we explored alternative sources for collecting parallel sentence pairs. In this study, we utilized a crawler to collect Japanese song lists from NetEase Cloud Music. We specifically targeted songs with human translations available as lyrics. After filtering and cleaning the data, we successfully collected about 700000 sentence pairs and constructed the WCC-JCL (Crawling a corpus of Japanese and Chinese song lyrics from the internet) corpus. We conducted extensive evaluations and experiments to validate the corpus’s effectiveness. The corpus WCC-JCL we crawled is open for free download by researchers, but only for research purposes.