Collecting Chinese-Vietnamese Texts From Bilingual Websites

Minh Trinh, Phuoc Dinh Tran, Nhung Tran · 2018

A monolingual-bilingual corpora are extremely necessary for natural language processing, especially for machine translation. In this paper, we propose a method to automatically collect bilingual Chinese-Vietnamese documents from bilingual Chinese-Vietnamese websites. These bilingual documents are the premise for extracting bilingual sentence pairs in our next research works. Our collection system was conducted on 10 Vietnamese-Chinese bilingual websites and initially gave encouraging results. This system can be deployed to collect automatically for other language pairs. less diversified.

Read the paper · More papers on PaperTik